Model details
magpie-tts-zeroshot
Magpie TTS Zeroshot is part of NVIDIA's Riva family of speech synthesis models, framed on the official NIM model card as a zero-shot text-to-speech system that produces expressive and engaging voice output from a short reference audio sample. The Overview section on build.nvidia.com explicitly positions it as a generative model intended to be paired with a companion neural audio codec in a two-stage pipeline, and the modelcard's category and filter labels tie it directly into the NVIDIA Riva TTS lineup, with an "Apply for Access" link pointing to the riva-tts-zeroshot-models page and an API reference under docs.nvidia.com/nim/riva/tts. The hosted page is marketed as a Free Endpoint accelerated by DGX Cloud, signaling NVIDIA's intent to make the model readily trialable rather than gated behind a sales motion, and the page is ready for commercial use under NVIDIA's foundation-model terms.
In practice the model is aimed at developers who need controllable voice, narration, and audio delivery without recording a full custom voice dataset: a brief sample of a target speaker is enough to drive generation. The NIM modelcard headline of "Expressive and engaging text-to-speech, generated from a short audio sample" describes the core zero-shot capability, while secondary directory coverage from Enterprise DNA echoes the same positioning as a speech generation model for controllable voice and narration, callable via the model identifier used here. Because the model is delivered as an NVIDIA NIM with a documented Riva TTS API surface, it fits naturally into existing Riva-based speech stacks, audio production pipelines, and prototyping workflows where consistent, brand-aligned voice output matters more than open-source flexibility.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/magpie-tts-zeroshot
- Release date
- May 22, 2025
- Last updated
- Jun 12, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 4,096 tokens
- Context window
- 0 tokens