Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

MiniMax M1

MiniMax M1 stands apart as the world's first open-weight hybrid-attention reasoning model, purpose-built to tackle complex tasks that demand both step-by-step reasoning and the ability to process extremely long inputs without drowning in computational costs. At its core lies a Mixture-of-Experts architecture housing 456 billion total parameters, though only about 45.9 billion activate for each token through sparse routing—effectively assembling a dynamic team of specialized experts that handle only the most relevant computations. This MoE foundation pairs with a novel Lightning Attention mechanism, which introduces linear-time attention optimized specifically for long sequences, dramatically cutting the compute overhead that typically scales quadratically with context length. The result is a decoder-only architecture that keeps inference efficient even when working with inputs that stretch into the hundreds of thousands of tokens.

The model's lineage traces back to a research paper focused on scaling test-time compute efficiently through Lightning Attention, reflecting a deliberate push toward practical reasoning under extended contexts. Rather than simply scaling pre-training blindly, the architecture emphasizes using test-time compute wisely—allowing the model to deliberate longer on difficult problems while maintaining speed on simpler ones. This design philosophy makes MiniMax M1 well-suited for applications like multi-turn customer support bots, code generation tools working with large repositories, and text analysis pipelines that need to extract meaning from lengthy documents. The combination of open-weight availability with a reasoning-optimized architecture positions it as a model that developers and researchers can both inspect and deploy flexibly across long-context use cases.

OpenRouterminimax/minimax-m1minimax

Quick Info

Powered by
Provider
OpenRouter
Model key
minimax/minimax-m1
Release date
Jun 17, 2025
Last updated
Jun 17, 2025
Knowledge cutoff
2024-06-30
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$2.20

Limits

Output tokens
40,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare MiniMax M1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiniMax M1

Qiniu

CoverageRelease Notes

A third-party release-tracker article dated May 19, 2026 documented that MiniMax-M1 has been superseded in MiniMax's current M-series lineup, with M3 (released June 1, 2026), M2.7, M2.5, M2.1, M2, and M2-her now listed in the API documentation rather than M1. The article treats M1 API availability and pricing as histor Despite that supersession, the article notes that MiniMax-M1 remains useful as an open-weight long-context reasoning model because its 1,000,000-token context window and open-weight release continue to attract long-document, code-reasoning, math, and agent-workload users. It frames M1 as a foundational release that est

Videos about MiniMax M1

More models around MiniMax M1