NovitaAI
Xiaomi's open-source large model MiMo-V2-Flash API has launched a payment function and is about to start a paid model. Users can log in to their accounts to obt
Model details
MiMo-V2-Flash is a Mixture-of-Experts language model built to balance massive scale with efficient execution. It features 309 billion total parameters while utilizing only 15 billion active parameters per forward pass, a design choice aimed at delivering high-speed reasoning. The architecture incorporates a hybrid attention mechanism that interleaves sliding window attention with global attention, which helps reduce the KV cache footprint significantly. By integrating Multi-Token Prediction, the model achieves faster inference speeds, making it well-suited for agentic workflows that require both depth and responsiveness.
The model was pre-trained on 27 trillion tokens and utilizes a Multi-Teacher On-Policy Distillation paradigm to refine its capabilities. This post-training method allows the model to learn from domain-specialized teachers, effectively absorbing expert knowledge to improve performance on complex tasks. With its ability to handle extended context lengths, the model is positioned to compete with frontier-level systems in coding and reasoning benchmarks. Its open-source nature and technical design make it a versatile tool for developers looking to integrate high-capacity, efficient AI into their own infrastructure.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
NovitaAI
Xiaomi's open-source large model MiMo-V2-Flash API has launched a payment function and is about to start a paid model. Users can log in to their accounts to obt
NovitaAI
Hi! I'm Niels, part of the community science team at Hugging Face.
NovitaAI
Xiaomi announced that the free public testing period of its self-developed large model MiMo-V2-Flash has been extended by 20 days, until January 20, 2026. The m
NovitaAI
Xiaomi unveils MiMo-V2-Flash open-source AI model, signaling strategic AI push.