Model details
Grok Voice STT 1.0
Grok Voice STT 1.0 is an audio-in, text-out speech recognition model from the Grok voice family, positioned as a transcription-focused counterpart to the broader Grok conversational lineup. Its appearance alongside Grok Voice TTS 1.0 in shared model catalogs suggests it was designed to handle inbound audio streams and convert spoken content into usable text, rather than to drive open-ended dialogue. The model name itself communicates the intended role: it is a speech-to-text specialist meant to complement the text-to-speech sibling for voice pipelines.
As a newly cataloged entry, Grok Voice STT 1.0 has not yet accumulated observable usage, efficiency, or peer-comparison signals, which is consistent with a fresh release. It fits naturally into voice agent stacks, media transcription pipelines, and accessibility tools where audio content needs to be turned into structured text quickly. Teams evaluating it should treat it as a specialized transcription module within the Grok ecosystem rather than a general-purpose chat model, pairing it with the matching TTS model when building full conversational voice experiences.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- x-ai/grok-voice-stt-1.0
- Release date
- Aug 4, 2026
- Last updated
- Aug 4, 2026
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 15,000 tokens
- Context window
- 15,000 tokens
Latest news about Grok Voice STT 1.0
No articles yet. Fetch the latest news to show it here.