GPT-6 Astra is OpenAI's flagship model for end-to-end digital work, designed around the idea of executing tasks across tools rather than only answering questions. It is OpenAI's largest training run to date, run on more than 100,000 GPUs at the Stargate site in Texas, and is the first release in which earlier OpenAI models helped supervise training of the new one. The model couples a roughly 1.05-million-token that quick-info value window with up to the cataloged API limit output tokens, multimodal text and image input, and configurable reasoning effort, so it can keep large codebases, document sets, or long agent histories in working memory and continue through more steps before asking for direction.
Astra posts OpenAI's strongest published scores on abstract reasoning, mathematics, and cybersecurity while pushing hard on computer-use and professional-work benchmarks. Reported figures include 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4 v2, 100% on ExploitBench, 72.6% on OSWorld 2.0 (in roughly 40 minutes versus the listed price for the prior model), and 41.4% on AutomationBench, with broad gains on Terminal-Bench, ScreenSpot-Pro, SRE-Bench, and long-context retrieval. Independent evaluators give a more nuanced picture, with the Artificial Analysis index favoring Claude and Meta's Muse Spark 1.3 edging Astra on agentic coding, so the model fits best as a specialist for computer use, long-running agent workflows, research, coding at scale, and tool-heavy enterprise tasks rather than as a universal leader on every benchmark.