Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

Grok Voice TTS 1.0

Grok Voice TTS 1.0 is SpaceXAI's text-to-speech model designed to turn written input into natural-sounding spoken audio. It handles more than twenty languages with automatic language detection, removing the need for explicit language tagging, and ships with five preset voices — Eve, Ara, Rex, Sal, and Leo — that cover a range of tones for different use cases such as narration, conversational agents, or character-driven content. Developers can shape delivery directly inside the prompt through inline speech tags that influence pauses, emphasis, pitch, speed, and vocal style, which makes the model flexible for both plain read-aloud and expressive scenarios without external post-processing.

On the output side, the model produces audio in MP3, WAV, PCM, μ-law, and A-law formats across a broad set of sample rates, letting integrators match the response to web playback, telephony pipelines, or embedded audio stacks. Per-request character limits align with a context-style budget suited to short scripts, dialogue turns, and segmented narration rather than full-document synthesis. Practical strengths highlighted in independent evaluation runs are fast turnaround on short utterances and reliable handling of speech-tag cues, while weaknesses appear on inputs that mix numbers, units, and unusual punctuation, where accuracy drops. The model fits best for product teams that need multilingual, voice-controlled synthesis with direct format flexibility and expressive control from a single API call.

ZenMuxx-ai/grok-voice-tts-1.0

Quick Info

Powered by
Provider
ZenMux
Model key
x-ai/grok-voice-tts-1.0
Release date
Jul 31, 2026
Last updated
Jul 31, 2026
Input modalities
Output modalities
Capabilities
Base catalog fields only

Limits

Output tokens
15,000 tokens
Context window
15,000 tokens

Latest news about Grok Voice TTS 1.0

ZenMux

CoverageBenchmark

Grok Voice TTS 1.0, attributed to SpaceXAI (xAI), is a text-to-speech model that converts text into spoken audio across 20+ languages with automatic language detection, offering five built-in voices (Eve, Ara, Rex, Sal, Leo) covering a range of tones. Per the OpenRouter model page, output is available in MP3, WAV, PCM, On the Artificial Analysis SpaceXAI TTS Elo benchmark, Grok Voice TTS 1.0 is reported with a score of 1,128 and a mean correctness of 91% across 12 prompts, ranking 2nd of 16 on correctness, 5th on cost (cheapest first), and 5th on speed (fastest first). Sample prompt-level results on the OpenRouter page show several T

Videos about Grok Voice TTS 1.0