Grok Voice TTS 1.0 is SpaceXAI's text-to-speech model designed to turn written input into natural-sounding spoken audio. It handles more than twenty languages with automatic language detection, removing the need for explicit language tagging, and ships with five preset voices — Eve, Ara, Rex, Sal, and Leo — that cover a range of tones for different use cases such as narration, conversational agents, or character-driven content. Developers can shape delivery directly inside the prompt through inline speech tags that influence pauses, emphasis, pitch, speed, and vocal style, which makes the model flexible for both plain read-aloud and expressive scenarios without external post-processing.
On the output side, the model produces audio in MP3, WAV, PCM, μ-law, and A-law formats across a broad set of sample rates, letting integrators match the response to web playback, telephony pipelines, or embedded audio stacks. Per-request character limits align with a context-style budget suited to short scripts, dialogue turns, and segmented narration rather than full-document synthesis. Practical strengths highlighted in independent evaluation runs are fast turnaround on short utterances and reliable handling of speech-tag cues, while weaknesses appear on inputs that mix numbers, units, and unusual punctuation, where accuracy drops. The model fits best for product teams that need multilingual, voice-controlled synthesis with direct format flexibility and expressive control from a single API call.