Currently listed through:
Model details
MAI-Code-1.1-Flash
MAI-Code-1.1-Flash is positioned as a small-tier developer assistant that builds on its predecessor with native image understanding, sharper coding quality, and tighter instruction following and tool use. Under the hood, it is a transformer that leans on self-attention routed through sparse Mixture-of-Experts layers, with 138 billion total parameters but only about 5 billion active per pass, a design that targets low latency and low serving cost while still handling complex agentic coding tasks such as working in real repositories, answering repository-level questions, refactoring, and tool-driven developer flows. The model card frames it as a text-to-text and image-to-text coding model intended for everyday developer work, with a 256,000-token context window that leaves room for sizable codebases and long agent traces.
Practically, the model's fit is in lightweight coding workflows where teams want capable behavior without paying for a heavyweight model: it is offered across the Copilot surface area, from the CLI and cloud agent to IDEs and mobile, and is exposed to free and student users through auto selection while paid SKUs can also pick it directly. Advances in model and serving efficiency translated into a 73 percent reduction in list price versus the prior generation, and annual Copilot subscribers use it at a 0.25× premium request multiplier, making it a cost-effective default for vision-aware coding help in Copilot Chat, coding agents, and developer tools.
Quick Info
Powered by- Provider
- GitHub Copilot
- Model key
- mai-code-1.1-flash
- Release date
- Aug 11, 2026
- Last updated
- Aug 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $1.20
Limits
- Input tokens
- 128,000 tokens
- Output tokens
- 128,000 tokens
- Context window
- 256,000 tokens