GPT-5.5 marks the first complete architectural rebuild since GPT-4.5, carrying the internal codename "Spud." Unlike the incremental updates that characterized every model released between those versions, this release represents a ground-up reconstruction designed from the start for natively omnimodal operation—handling text, images, audio, and video through a single unified system rather than stitching separate modality models together. The model was co-designed with NVIDIA's GB200 and GB300 NVL72 rack-scale systems, a partnership that enabled it to achieve GPT-5.4-level per-token latency despite significantly expanded capabilities. Larger models typically sacrifice speed; GPT-5.5 defies that tradeoff through tight hardware-software co-optimization.
Beyond the model itself, GPT-5.5 and its companion system Codex rewrote OpenAI's own serving infrastructure before launch, analyzing weeks of production traffic to optimize performance end-to-end. The release also introduced GPT-5.5 Instant as the new default model for ChatGPT, bringing these improvements to everyday users while specifically reducing hallucination rates in sensitive domains including law, medicine, and finance—areas where accuracy matters most. The combination of architectural redesign, hardware co-design, and infrastructure-level improvements positions GPT-5.5 as a platform-level advancement rather than a simple capability bump, extending its practical fit beyond general chat into agentic workflows and high-stakes professional applications.