LLM Gateway
Grok 4.20 leads BridgeBench reasoning benchmark, surpassing OpenAI's GPT-5.4 and other major competitors, as announced by WesRoth on April 15.
Model details
Grok 4.20 Reasoning represents xAI's flagship push into advanced reasoning-capable language models, designed with dual-variant flexibility that lets developers choose between reasoning and non-reasoning modes depending on task demands. The model was engineered with a clear emphasis on reliability: by targeting the lowest hallucination rate available and enforcing strict prompt adherence, Grok 4.20 seeks to deliver consistent, trustworthy outputs that stay true to user intent. Its agentic tool-calling capabilities position it for workflows that require connecting language understanding to external systems and actions, while its expansive context window enables deep analysis across lengthy documents or multi-turn conversations without losing thread.
The model's competitive positioning comes from recent benchmark results showing it atop BridgeBench, a dedicated reasoning evaluation, where it surpassed GPT-5.4 and other leading competitors. This achievement underscores the effectiveness of xAI's approach to cultivating strong reasoning behavior. Grok 4.20's combination of speed, precision, and tool integration makes it well-suited for enterprise deployments and research applications where accuracy and operational reliability matter most. The focus on reducing fabrication while maintaining prompt fidelity reflects a maturation of the Grok family beyond earlier iterations, targeting users who need dependable AI assistance in complex, high-stakes environments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
LLM Gateway
Grok 4.20 leads BridgeBench reasoning benchmark, surpassing OpenAI's GPT-5.4 and other major competitors, as announced by WesRoth on April 15.