Grok Voice Think Fast 2.0 is a speech-to-speech voice model announced on July 29, 2026 as the next generation of SpaceXAI's conversational voice line. Rather than chaining separate speech-to-text and text-to-speech components, the model accepts audio and streams audio back directly over a WebSocket connection, letting it reason about what it hears while it is still talking. This unified architecture is designed to remove the awkward pause before the first reply and to keep the agent responsive even when function calls are running in the middle of a turn, which is a common pain point in production voice stacks.
In practical use, the model behaves like a true full-duplex conversational partner: it listens while it speaks instead of waiting for strict turn boundaries, handles mid-sentence interruptions gracefully, and can recover after a dropped connection. It also fires tool and function calls early in a turn so that downstream actions such as looking up an order or updating a delivery instruction begin while audio is still streaming, which tightens the loop for real-world customer support and transactional voice agents. Coverage positions it as faster to first audio and more accurate on transcription under noisy conditions than its predecessor, making it a fit for teams building always-on phone, in-car, or kiosk assistants where latency and robustness matter more than open customization.