Currently listed through these providers:
Model details
nemotron-voicechat
The Nemotron VoiceChat model represents a specialized speech decoder developed within NVIDIA's NeMo framework, targeting real-time voice interaction workflows. Its architecture is engineered for duplex speech-to-speech systems, meaning it can handle bidirectional conversational flows essential for natural dialogue. The model emphasizes low-latency performance and multilingual deployment capabilities, making it suitable for conversational AI applications that require responsive, human-like interactions across different languages.
Development of this model involved contributions from researchers working directly on the NeMo repository, with specific enhancements to text-to-speech model catalogs and audio processing pipelines. The implementation supports half-precision inference for computational efficiency, while design choices prioritize reproducible training, scalable deployment, and integration with broader conversational AI systems. Its positioning as an open-weight model enables organizations to fine-tune and deploy the speech decoder within their own infrastructure for custom voice applications.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/nemotron-voicechat
- Release date
- Mar 16, 2026
- Last updated
- Mar 16, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 8,192 tokens
- Context window
- 128,000 tokens