LLM Gateway
MiniMax M2.1 Lightning is a faster variant of M2.1 with lower latency for coding tasks.
Model details
MiniMax M2.1 Lightning is a performance-oriented member of MiniMax's M2 series, engineered as a faster alternative to the standard M2.1 with a specific emphasis on reducing latency for coding and agentic workflows. The model builds on the foundation of its predecessor while introducing optimizations that make it particularly well-suited for applications where response speed directly impacts user experience or task efficiency. Its open-weight status means developers can inspect, fine-tune, and deploy it in self-hosted environments, giving teams more control over their inference infrastructure.
The Lightning variant achieves approximately 100 tokens per second output speed, a characteristic that positions it as a practical choice for latency-sensitive scenarios such as interactive coding assistants, real-time agents, and streaming applications. While the sources describe the model as accelerated and optimized for throughput, they do not disclose the specific architectural modifications or training recipes responsible for these gains. It retains core capabilities from the M2 family, including function calling and reasoning support, and its million-token context window enables it to handle extended documents, codebases, or multi-turn conversations without frequent resets.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
LLM Gateway
MiniMax M2.1 Lightning is a faster variant of M2.1 with lower latency for coding tasks.