MiniMax Token Plan (minimax.cn)
The official MiniMax-M2 series technical report (arXiv:2605.26494) introduces the M2 family of Mixture-of-Experts language models built on the principle that minimal activated parameters can deliver high real-world intelligence. The flagship M2 is configured with 229.9B total parameters and only 9.8B activated per toke According to the paper, M2 is trained on large-scale verifiable trajectories spanning agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward. Forge is described as a scalable agent-native RL system that uses windowed-FIFO scheduling, prefix-tree merging, and a clean t