Grok 4 is xAI's flagship reasoning model, positioned at launch as the most capable entry in the Grok family and built on lessons from Grok the listed price Reasoning. xAI reports that it scaled reinforcement learning training using its Colossus 200,000-GPU cluster, achieving roughly 6x compute efficiency gains over prior runs and extending verifiable training data well beyond math and coding into many more domains. The model was introduced with native tool use and real-time search integration, and a higher-tier Grok 4 Heavy variant enables multiple agents to brainstorm together while invoking tools such as code execution and web search for harder problems.
In independent benchmarking, Grok 4 posted a leading score on Humanity's Last Exam, a 2,500-question multimodal benchmark from the Center for AI Safety and Scale AI, outperforming competing frontier models by a wide margin, and it is among the systems evaluated in a 2026 Scientific Reports pediatric dentistry benchmarking study. Distribution expanded beyond consumer subscriptions and the xAI API into enterprise channels, with Microsoft announcing Grok 4 availability on Azure AI Foundry in early 2026. Practically, Grok 4 fits well for research-style reasoning, agentic workflows that benefit from tool-calling and live search, and domain question-answering where strong academic and multi-modal performance matter more than open-weight flexibility.