# NVIDIA VoiceChat-11B
NVIDIA-NemotronLabs-VoiceChat-11B is an open-weights voice model NVIDIA put on Hugging Face on 3 August 2026. It's one speech-to-speech model with 11 billion parameters: audio in, audio out, with no hand-off between separate speech recognition, language and voice models.
It's **full-duplex**, so it listens while it speaks and you can cut it off mid-sentence. It can also **call tools** while the conversation keeps going. NVIDIA calls it "the first open full-duplex model to support tool calling". I can't verify the "first", but I don't know of an earlier open model that does both.
## One model instead of three
Most voice agents chain three models: [[Speech-to-Text (STT)]] transcribes you, an [[Large Language Models (LLMs)|LLM]] writes the answer and [[Text-to-Speech (TTS)]] reads it out. Each step waits for the previous one, every hand-off adds latency, and the agent can't hear you while it talks, so interrupting it is clumsy. Doing it all in one network brings turn-taking close to human timing and makes barge-in (talking over the agent) work.
## How it's built
A hybrid Mamba/Transformer (Mamba is a state-space architecture that handles long sequences more cheaply than attention), assembled from NVIDIA parts:
- a Fast Conformer speech encoder taken from Nemotron-Speech-Streaming
- Nemotron Nano v2 (9B) as the language backbone, from the [[NVIDIA Nemotron]] family
- NVIDIA's own TTS decoder and audio codec for the voice
- a separate output stream for function calls, next to the speech
Inputs are text and 16 kHz audio. Outputs are text, 22.05 kHz audio and a transcript of what the user said. While a tool runs, it can play a predefined "on hold" message so the line doesn't go quiet.
Training used about 550,000 hours of audio: real speech (Fisher, LibriVox, VCTK, studio recordings) plus synthetic speech made by running TTS over text datasets such as Ultrachat and NVIDIA's Nemotron data.
## Numbers
| Metric | Result |
|---|---|
| Turn-taking latency (Full-Duplex-Bench 1.0) | ~450 ms |
| User interruption latency (Full-Duplex-Bench 1.0) | ~480 ms |
| VoiceBench | 2nd among open full-duplex models |
| BFCL-v3 tool calling, average (AU Harness) | 56.1% |
| Tool selection (Full-Duplex-Bench v3) | 82.5% |
| Parallel multiple tools | 27.5% |
| Argument accuracy | 44.2% |
For agents, the bottom three rows count most: it picks the right tool 82.5% of the time, gets the arguments right less than half the time, and mostly fails at parallel calls. Good enough for a "look up my order" agent. For multi-step workflows, I'd want a text model checking its work.
## Running it
Linux only, on NVIDIA data center or workstation GPUs (A100, H100, H200, B100, B200, RTX 6000), with [[vLLM]] for inference. Weights are on [[HuggingFace]]. It's built for servers, so forget the laptop.
## License
The weights ship under the OpenMDW License Agreement v1.1. HackerNoon's write-up calls it a research model, so read the license before you build a product on it.
## How it compares
Kyutai's Moshi was the earlier open full-duplex model, without tool calling. Sesame has its own conversational voice models. On the closed side there's OpenAI's [[GPT-Realtime-2]] through the [[OpenAI Realtime API]], and [[ElevenLabs]] sells cascaded voice agents on top of its TTS.
My take: a separate tool-call channel is how I'd build [[AI Agents]] that talk, because the voice keeps going while a function runs instead of the call freezing. With 44% argument accuracy, though, I'd use it as the conversational front-end and hand anything with real arguments to a stronger text model.
## References
- NVIDIA model card on Hugging Face: https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
- aimodels44, "NVIDIA VoiceChat-11B brings full-duplex AI speech to real-time agents", HackerNoon, 2026-08-06: https://hackernoon.com/nvidia-voicechat-11b-brings-full-duplex-ai-speech-to-real-time-agents
## Related
- [[NVIDIA Nemotron]]
- [[LLM Tool Calling]]
- [[AI Tool Use]]
- [[AI Agents]]
- [[AI Open Weight Models]]
- [[Speech-to-Text (STT)]]
- [[Text-to-Speech (TTS)]]
- [[GPT-Realtime-2]]
- [[OpenAI Realtime API]]
- [[ElevenLabs]]
- [[vLLM]]
- [[HuggingFace]]
- [[Gemini 3.1 Flash Live]]
- [[Open-LLM-VTuber]]