VoiceChat Preserves Turn Taking With Native Tool Calls
A streaming speech stack separates response text, function calls, transcription, and synthesis, reaching 82.5% tool-selection F1 while maintaining interruption handling.
Underlying Paper
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100\% takeover following user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes its response after user backchannels in 93\% of cases. NemotronLabs VoiceChat obtains a 55.1 normalized average on VoiceBench and, on Full-Duplex-Bench 3.0 (FDB 3.0), achieves 82.5\% tool-selection F1, while argument accuracy and end-to-end tool execution remain areas for improvement. These results demonstrate that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.
Voice agents usually assemble speech recognition, a language model, text-to-speech, and tool execution as separate services. That division makes timing difficult: an agent must know when a short user acknowledgement is only a backchannel, when an interruption should take over, and how to keep speaking while a tool is running. NemotronLabs VoiceChat puts those behaviors into one streaming speech-to-speech model, with an auxiliary transcription path rather than a text-first conversational pipeline.
Core Contribution
The paper's central design choice is to make conversational outputs parallel rather than sequential. A streaming speech encoder feeds a decoder-only language model that produces agent-response text and structured function calls through distinct heads, while an RNN-T branch incrementally transcribes the user and a streaming TTS decoder produces speech. The split matters because text intended for the user and a tool invocation have different temporal requirements: the former can be delayed slightly to avoid talking over new audio, while the latter must retain its position in the dialogue.
The authors frame this as a full-duplex system rather than a voice interface with endpoint detection. During supervised fine-tuning, they create early interruptions, insert recorded backchannels, and shift agent-text targets by two 80 ms frames. Function-call targets are not shifted. This is a concrete attempt to teach the model that an “uh-huh” should not terminate an answer, while a real interruption should.
Technical Approach
The architecture diagram shows the streaming path from encoded audio into the language-model backbone, alongside specialized text and function-call channels.
At inference, the FastConformer encoder consumes 16 kHz audio with chunk-aware caching. Its projected audio embeddings go to the language model, while the raw encoder features feed the RNN-T side channel. The TTS decoder autoregressively emits one 31-codebook audio index per 80 ms model step, then a causal PyTorch codec decodes those indices into 22.05 kHz waveform samples. The runtime uses cached state, CUDA Graphs, a custom LLM fork, and Triton serving to avoid treating every audio chunk as a new request.
Tool use is handled on its own channel. Once the model emits a tool-call start marker, the runtime asynchronously completes the call and plays a predefined acknowledgement while execution is pending. Returned results are inserted into the function-call channel and decoder context before response generation resumes. The tool-channel schematic captures that separation between agent text and structured calls.
The training recipe is similarly targeted. CPT is mostly speech-text pretraining synthesized from text passages with alternating user and agent roles. SFT adds conversational, instruction-following, tool-calling, and safety data. The SFT sampling mixture assigns normalized shares of about 46.8% speech-text pretraining, 23.8% conversational behavior, 26.0% tool calling, and 3.4% safety. Tool data is generated with a multi-agent pipeline because raw tool traces contain URLs, Markdown, and code that do not naturally map to spoken turns.
Results and Analysis
The full-duplex results are the paper's strongest evidence. On Full-Duplex-Bench 1.0, the model reports the lowest pause-handling takeover rates among the evaluated open-weight systems, 100% takeover after user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes after user backchannels in 93% of cases. These are behavior-specific tests rather than a generic speech score, so they support the claim that the augmentation and output timing are doing useful work.
The system is also evaluated beyond turn management. It reaches a 55.1 normalized average on VoiceBench and 82.5% F1 for tool selection on FDB 3.0. The latter is a meaningful distinction: selecting a tool is not the same as providing correct arguments or completing the full execution path, and the authors explicitly identify both as weaker areas.
Speech generation remains competitive but is not the best conventional TTS result. On an unseen-speaker first turn, VoiceChat-TTS records 2.00% WER and 4.380 SQuIM-MOS, compared with Audio Flamingo 3-Chat's 4.51% WER and 3.600 SQuIM-MOS. Across four turns, unseen-speaker WER changes from 2.00% to 2.20%, while speaker similarity falls from 0.757 to 0.685. That trade-off is acceptable for an interactive model that remains active through overlapping speech, but it shows that persistent synthesis can drift over longer exchanges.
For streaming transcription, increasing the chunk size from 80 ms to 160 ms reduces average WER from 9.02% to 8.28% across eight OpenASR datasets. With four concurrent streams on one NVIDIA H100 PCIe GPU, the reported p95 latency is 118 ms per 160 ms audio chunk, or 1.36× real-time throughput. The measurements establish a usable real-time operating point, though they do not show latency under tool execution or on lower-end deployment hardware.
Evidence Box
strongKey Claims
- •Unified streaming speech, transcription, reasoning, synthesis, and tool calling
- •Full-duplex behavior that distinguishes interruptions from backchannels
- •Structured function calls generated on a dedicated output channel
- •Low-latency concurrent streaming inference
Key Results
- •100% interruption takeover and 4.33/5 post-interruption quality on Full-Duplex-Bench 1.0
- •93% response resumption after backchannels on Full-Duplex-Bench 1.5
- •82.5% tool-selection F1 on Full-Duplex-Bench 3.0
- •8.28% average ASR WER at 160 ms chunks vs. 9.02% at 80 ms
Limitations & Caveats
- •Tool argument accuracy and end-to-end execution remain weaker than 82.5% tool-selection F1
- •Barge-in is unavailable while a pending tool call is being executed
- •Inference latency reported only on an NVIDIA H100 PCIe with four concurrent streams
- •Unseen-speaker similarity falls from 0.757 to 0.685 across four TTS turns