# NVIDIA VoiceChat-11B NVIDIA-NemotronLabs-VoiceChat-11B is an open-weights voice model NVIDIA put on Hugging Face on 3 August 2026. It's one speech-to-speech model with 11 billion parameters: audio in, audio out, with no hand-off between separate speech recognition, language and voice models. It's **full-duplex**, so it listens while it speaks and you can cut it off mid-sentence. It can also **call tools** while the conversation keeps going. NVIDIA calls it "the first open full-duplex model to support tool calling". I can't verify the "first", but I don't know of an earlier open model that does both. ## One model instead of three Most voice agents chain three models: [[Speech-to-Text (STT)]] transcribes you, an [[Large Language Models (LLMs)|LLM]] writes the answer and [[Text-to-Speech (TTS)]] reads it out. Each step waits for the previous one, every hand-off adds latency, and the agent can't hear you while it talks, so interrupting it is clumsy. Doing it all in one network brings turn-taking close to human timing and makes barge-in (talking over the agent) work. ## How it's built A hybrid Mamba/Transformer (Mamba is a state-space architecture that handles long sequences more cheaply than attention), assembled from NVIDIA parts: - a Fast Conformer speech encoder taken from Nemotron-Speech-Streaming - Nemotron Nano v2 (9B) as the language backbone, from the [[NVIDIA Nemotron]] family - NVIDIA's own TTS decoder and audio codec for the voice - a separate output stream for function calls, next to the speech Inputs are text and 16 kHz audio. Outputs are text, 22.05 kHz audio and a transcript of what the user said. While a tool runs, it can play a predefined "on hold" message so the line doesn't go quiet. Training used about 550,000 hours of audio: real speech (Fisher, LibriVox, VCTK, studio recordings) plus synthetic speech made by running TTS over text datasets such as Ultrachat and NVIDIA's Nemotron data. ## Numbers | Metric | Result | |---|---| | Turn-taking latency (Full-Duplex-Bench 1.0) | ~450 ms | | User interruption latency (Full-Duplex-Bench 1.0) | ~480 ms | | VoiceBench | 2nd among open full-duplex models | | BFCL-v3 tool calling, average (AU Harness) | 56.1% | | Tool selection (Full-Duplex-Bench v3) | 82.5% | | Parallel multiple tools | 27.5% | | Argument accuracy | 44.2% | For agents, the bottom three rows count most: it picks the right tool 82.5% of the time, gets the arguments right less than half the time, and mostly fails at parallel calls. Good enough for a "look up my order" agent. For multi-step workflows, I'd want a text model checking its work. ## Running it Linux only, on NVIDIA data center or workstation GPUs (A100, H100, H200, B100, B200, RTX 6000), with [[vLLM]] for inference. Weights are on [[HuggingFace]]. It's built for servers, so forget the laptop. ## License The weights ship under the OpenMDW License Agreement v1.1. HackerNoon's write-up calls it a research model, so read the license before you build a product on it. ## How it compares Kyutai's Moshi was the earlier open full-duplex model, without tool calling. Sesame has its own conversational voice models. On the closed side there's OpenAI's [[GPT-Realtime-2]] through the [[OpenAI Realtime API]], and [[ElevenLabs]] sells cascaded voice agents on top of its TTS. My take: a separate tool-call channel is how I'd build [[AI Agents]] that talk, because the voice keeps going while a function runs instead of the call freezing. With 44% argument accuracy, though, I'd use it as the conversational front-end and hand anything with real arguments to a stronger text model. ## References - NVIDIA model card on Hugging Face: https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B - aimodels44, "NVIDIA VoiceChat-11B brings full-duplex AI speech to real-time agents", HackerNoon, 2026-08-06: https://hackernoon.com/nvidia-voicechat-11b-brings-full-duplex-ai-speech-to-real-time-agents ## Related - [[NVIDIA Nemotron]] - [[LLM Tool Calling]] - [[AI Tool Use]] - [[AI Agents]] - [[AI Open Weight Models]] - [[Speech-to-Text (STT)]] - [[Text-to-Speech (TTS)]] - [[GPT-Realtime-2]] - [[OpenAI Realtime API]] - [[ElevenLabs]] - [[vLLM]] - [[HuggingFace]] - [[Gemini 3.1 Flash Live]] - [[Open-LLM-VTuber]]