OpenAI
Introducing GPT-5.3-Codex-Spark—our first real-time coding model. 15x faster generation, 128k context, now in research preview for ChatGPT Pro users.
Model details
GPT-5.3 Codex Spark is positioned as OpenAI's first real-time coding model, designed for interactive development workflows where responsiveness matters as much as raw capability. It belongs to the gpt-codex-spark family and serves as a smaller, speed-tuned variant of GPT-5.3-Codex, optimized for targeted edits, logic adjustments, and interface refinements rather than large-scale generation tasks. The model is released as a research preview and marks OpenAI's first production deployment running on Cerebras Wafer Scale Engine 3 hardware, stepping outside the company's traditional NVIDIA-based infrastructure.
In practice, GPT-5.3 Codex Spark delivers generation speeds that enable near-instant feedback during coding sessions, with reported sustained throughput exceeding 1,000 tokens per second. This makes it well-suited for scenarios such as making targeted code edits, refining application logic, and iterating on user interfaces in real time. Its text-only input design aligns with its focus on code-centric tasks, and developers can expect faster cycle times compared to larger Codex variants, making it a practical choice for interactive tooling and agentic coding assistants where latency is a primary concern.
OpenAI
Introducing GPT-5.3-Codex-Spark—our first real-time coding model. 15x faster generation, 128k context, now in research preview for ChatGPT Pro users.
OpenAI
OpenAI released GPT-5.3-Codex-Spark as a research preview, describing it as a smaller variant of GPT-5.3-Codex and its first model purpose-built for real-time coding inside Codex. The launch marks the first milestone in OpenAI's January-announced partnership with Cerebras, and the model is served on Cerebras hardware t At launch, Codex-Spark is text-only with a 128k context window and has its own rate limits that do not count against standard Codex limits, though OpenAI warned of temporary queuing under high demand. It is tuned for interactive use, making minimal targeted edits by default and skipping automatic test runs unless asked
OpenAI
In a major shift in its hardware strategy, OpenAI launched GPT-5.3-Codex-Spark, its first production AI model deployed on Cerebras wafer-scale chips rather than traditional Nvidia GPUs. The new model
OpenAI
OpenAI released the Codex-Spark coding model to help developers to make targeted edits, adjust logic, refine interfaces, and more.