Listen, speak, handle interruptions, and call tools in one live conversation. 🎙️ Introducing NVIDIA-NemotronLabs-VoiceChat-11B, NVIDIA’s end-to-end full-duplex model for real-time voice agents.
🤖 https://t.co/Nqkz5qgw1l
🏆 It ranks #2 among open full-duplex models on both VoiceBench and Full-Duplex-Bench 1.0.
⚡ Natural turn-taking responds in ~448 ms, while barge-in lets users interrupt the model with ~480 ms latency.
🛠️ The first open full-duplex model with live tool calling. One unified system handles speech understanding, generation, transcription, and tool execution while keeping the conversation flowing.
🧠 Built with a Fast Conformer, Nemotron Nano v2, and a TTS decoder, trained on ~550K hours of real and synthetic speech.
📜 Research use only. OpenMDW 1.1.