Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Grok Voice Think Fast 2.0

Grok Voice Think Fast 2.0 is a speech-to-speech voice model announced on July 29, 2026 as the next generation of SpaceXAI's conversational voice line. Rather than chaining separate speech-to-text and text-to-speech components, the model accepts audio and streams audio back directly over a WebSocket connection, letting it reason about what it hears while it is still talking. This unified architecture is designed to remove the awkward pause before the first reply and to keep the agent responsive even when function calls are running in the middle of a turn, which is a common pain point in production voice stacks.

In practical use, the model behaves like a true full-duplex conversational partner: it listens while it speaks instead of waiting for strict turn boundaries, handles mid-sentence interruptions gracefully, and can recover after a dropped connection. It also fires tool and function calls early in a turn so that downstream actions such as looking up an order or updating a delivery instruction begin while audio is still streaming, which tightens the loop for real-world customer support and transactional voice agents. Coverage positions it as faster to first audio and more accurate on transcription under noisy conditions than its predecessor, making it a fit for teams building always-on phone, in-car, or kiosk assistants where latency and robustness matter more than open customization.

Vercel AI Gatewayspacexai/grok-voice-think-fast-2.0grok

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
spacexai/grok-voice-think-fast-2.0
Release date
Jul 29, 2026
Last updated
Jul 29, 2026
Input modalities
Output modalities
Capabilities
Base catalog fields only

Limits

Output tokens
0 tokens
Context window
0 tokens

Latest news about Grok Voice Think Fast 2.0

Vercel AI Gateway

Coverage

xAI released Grok Voice Think Fast 2.0 on July 29, 2026, positioning it as a substantial upgrade over Think Fast 1.0 for enterprise voice AI workloads, according to this Enterprise DNA summary attributed to xAI. The model reportedly cuts time-to-first-audio from 1.25s to 0.70s, delivers 1.4x better transcription accura Pricing is set at $0.08 per minute of audio, and xAI cites A/B testing on Starlink showing higher sales conversion and support containment rates for the new model. The article frames Think Fast 2.0 as closing the gap with purpose-built voice AI leaders while maintaining competitive economics for call-center, field-serv

Vercel AI Gateway

CoverageRelease Notes

Grok Voice Think Fast 2.0, announced on July 29, 2026, is a speech-to-speech model that reasons in parallel with speech—taking audio in and producing audio out while thinking through the query mid-response. This parallelism enables it to be substantially smarter than other speech-to-speech models with no added latency, Headline numbers from the launch include an Artificial Analysis Speech-to-Speech Quality Index score of 82.9% (up from 75.7% in version 1.0), time to first audio of 0.70s (down from 1.25s), reasoning tokens at roughly 0.4× the volume of version 1.0 per response, and pricing at $0.08 per minute of audio. SpaceXAI's publ

Vercel AI Gateway

Coverage

SpaceXAI launched Grok Voice Think Fast 2.0 on July 29, 2026, positioning it as the successor to Grok Voice Think Fast 1.0 and their most intelligent voice model yet. The model is designed to reason in parallel with speech, making it substantially smarter than other speech-to-speech models with no impact on latency, ac Grok Voice Think Fast 2.0 is priced at $0.08 per minute of audio. SpaceXAI benchmarked it against Grok Voice Think Fast 1.0, OpenAI's GPT-Realtime-2.1, and Google's Gemini 3.1 Flash, claiming it outperformed the three on most tests. During A/B testing on Starlink, the company reported a significant increase in sales co

Vercel AI Gateway

CoverageAnalysis

Grok Voice Think Fast 2.0 represents an escalation in the multimodal LLM competition, shifting focus from raw reasoning power to raw inference speed for real-time voice applications. The model moves away from cascaded ASR-LLM-TTS pipelines toward an end-to-end multimodal design that works directly with audio tokens, ai The release positions Grok Voice Think Fast 2.0 against OpenAI's Advanced Voice and Google's Gemini Live in the push for natural, interruption-tolerant exchanges. Key technical demands include persistent low-latency streaming, clean handling of stuttering, cross-talk, and background noise, while maintaining competitive

Vercel AI Gateway

CoverageBenchmark

SpaceXAI shipped Grok Voice Think Fast 2.0 on July 29, 2026, with the model scoring 82.9% on Artificial Analysis' speech-to-speech benchmark—beating GPT-Realtime-2.1 at 79.1% and Gemini 3.1 Flash at 69.5%. The previous version, Think Fast 1.0, scored 75.7% on the same benchmark, reflecting a 7-point jump over the prior Key improvements center on latency (time to first audio dropped from 1.25 seconds to 0.70 seconds), transcription accuracy (1.5 to 2× better compared to Deepgram Nova 3 and ElevenLabs Scribe v2 on their evaluation set), and token efficiency (roughly 60% fewer reasoning tokens than version 1.0). The model supports more

Vercel AI Gateway

Coverage

On July 29, 2026, xAI (SpaceXAI) announced Grok Voice Think Fast 2.0 with better speech reasoning, 1.5–2× transcription word error rate improvement versus dedicated STT on their 24-language evaluation, 60% fewer reasoning tokens than version 1.0, and faster tool timing. The alias grok-voice-latest flips to 2.0 on Augus Technical highlights include an AA Speech-to-Speech overall score of 82.9% (up from 75.7% for 1.0), τ-voice Bench at 56.5% versus 52.1% for 1.0, and time to first audio dropping from 1.25 seconds to 0.70 seconds. The model is accessible via the wss://api.x.ai/v1/realtime endpoint with WebSocket/Realtime-compatible surf

Vercel AI Gateway

Coverage

Bleap's explainer profiles Grok Voice Think Fast 2.0 as xAI's low-latency, real-time speech-to-speech voice model launched on July 29, 2026, designed for developers building voice agents in customer service, sales, and telephone scenarios. Its architecture listens, reasons, and speaks simultaneously without waiting for Key reported specs include $0.08 per audio minute pricing, a 1.4x transcription accuracy improvement over Think Fast 1.0, gains in conversational behavior and tool use, and coverage alongside comparisons with OpenAI and Gemini voice offerings. The piece positions Think Fast 2.0 as xAI's fastest voice model in the Grok

Vercel AI Gateway

Coverage

The official SpaceXAI news index lists "Introducing Grok Voice Think Fast 2.0" dated July 29, 2026, described as "our most capable speech-to-speech voice model." The index entry confirms the model's official launch and branding as SpaceXAI's next-generation voice offering within the broader Grok model family. This official listing places Grok Voice Think Fast 2.0 alongside other recent SpaceXAI product releases including Grok 4.6, Grok Bot, Imagine Image 2.0, and Imagine Video 1.5, confirming its position in the company's 2026 model roadmap. The entry serves as the first-party confirmation of the model's existence and relea

Videos about Grok Voice Think Fast 2.0

More models around Grok Voice Think Fast 2.0