Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Qiniu logo

Model details

Mimo-V2-Flash

MiMo-V2-Flash is a Mixture-of-Experts language model engineered to balance high-level intelligence with operational efficiency. It utilizes a sparse architecture featuring 309 billion total parameters, with only 15 billion active parameters per inference, allowing it to maintain speed while handling demanding tasks. The model is built on a novel hybrid attention architecture that interleaves a 128-token sliding window with full attention in a 5:1 ratio. This design choice significantly reduces KV-cache storage requirements, making the model particularly effective for long-context applications and complex reasoning scenarios where performance and resource management are critical.

The model incorporates Multi-Token Prediction to enhance its generative capabilities, positioning it as a strong competitor in software engineering and scientific reasoning benchmarks. Its design lineage focuses on practical utility, showing top-tier performance in coding evaluations like SWE-bench and scientific assessments such as GPQA-Diamond. By optimizing the interaction between its 256 experts, the model achieves a lightweight footprint relative to its total parameter count, offering a versatile tool for developers and researchers who require a robust foundation for agentic workflows and general-purpose assistance.

Qiniumimo-v2-flashmimo

Quick Info

Powered by
Provider
Qiniu
Model key
mimo-v2-flash
Release date
Dec 16, 2025
Last updated
Feb 4, 2026
Knowledge cutoff
2024-12-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.30

Limits

Output tokens
256,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Mimo-V2-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mimo-V2-Flash

Qiniu

Coverage

At the 2025 Xiaomi Human-Car-Home Ecosystem Partner Conference, newly appointed MiMO large-model lead Luo Fuli officially unveiled MiMo-V2-Flash as Xiaomi's latest MoE (Mixture of Experts) large model, framed as the company's second step toward its Artificial General Intelligence goal. The article, published on Decembe Technically, MiMo-V2-Flash adopts a Hybrid Sliding Window Attention (SWA) architecture with an optimal window size of 128 and a fixed KV cache for infrastructure compatibility, delivering improved long-context reasoning compared to other linear attention variants. The model also incorporates Multi-Token Prediction (MTP

Videos about Mimo-V2-Flash

More models around Mimo-V2-Flash