Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
AMD logo

Model details

MiniCPM5-2B

MiniCPM5-2B is an efficient small-scale language model developed by OpenBMB, the open-source group behind the MiniCPM series, and released under the permissive Apache 2.0 license. It is a dense model with roughly 2.6 billion total parameters and a text-only interface, designed to bring capable reasoning into resource-constrained environments such as smartphones, PCs, and embedded devices. The model fits within OpenBMB's broader strategy of producing compact architectures that can serve as foundations for local intelligent agents rather than simple on-device question answering, with the dense design favoring a smaller memory footprint over reductions in active-parameter compute.

On the Artificial Analysis Intelligence Index v4.2, MiniCPM5-2B posts a score of 15, the highest of any open-weights model with fewer than 4 billion total parameters and a clear step ahead of comparably sized peers such as Granite 4.2 3B. It also demonstrates notable agentic strength for its class, leading sub-4B models on GDPval-AA v2 Elo and tying for first on τ³-Banking, while remaining competitive on AA-Briefcase. These results suggest a practical fit for developers who need a lightweight, open reasoning model capable of tool use, long-context handling, and agent-style workflows without the cost or footprint of larger frontier systems.

AMDMiniCPM5-2B

Quick Info

Powered by
Provider
AMD
Model key
MiniCPM5-2B
Release date
Sep 6, 2026
Last updated
Sep 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.124
Output token cost
$0.7425

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Latest news about MiniCPM5-2B

AMD

Coverage

Medium's independent analysis notes that OpenBMB released MiniCPM5-2B on September 7, 2026: a dense, 2.516-billion-parameter causal language model under Apache 2.0 with a native 131,072-token context window. The headline claim is that this 2B-class model is state-of-the-art among sub-4B open models, averaging 53.9 acro The article reports that when MiniCPM5-2B is wired into a real multi-turn agentic harness, its performance collapses by up to 3x relative to single-shot index scores, highlighting a gap between elite single-shot coding benchmarks and chaotic production software engineering realities. It frames the model as benchmark-sp

AMD

CoverageBenchmark

SitePoint's independent benchmark analysis reports that on Artificial Analysis's GDPval-AA v2, which scores models on real-world work tasks against a 1,000 human baseline, MiniCPM5-2B scores 831 and finishes first among open-weight models under 4B parameters, ahead of Ling 3.0 Tiny (718), Granite 4.2 8B (648), Gemma 4 On the composite Intelligence Index that GDPval-AA v2 feeds into, MiniCPM5-2B places third, behind substantially larger models including Ling 3.0 Tiny at roughly three times the parameters and Gemma 4 12B at several times larger. The article frames the dual result as evidence about how to read benchmark scores and capa

AMD

CoverageRelease Notes

DataNorth reports that OpenBMB released MiniCPM5-2B on 7 September 2026, an open dense model with 2.52 billion parameters under Apache 2.0 and a context window of 131,072 tokens. On OpenBMB's own 34-benchmark comparison table, MiniCPM5-2B averages 53.9 versus 51.1 for Qwen3.5-4B, which carries roughly twice the paramet The model targets hardware a developer controls (laptop, phone, single small GPU) rather than a rented data center, and uses the standard LlamaForCausalLM layout so vLLM, SGLang, Transformers, llama.cpp, Ollama, LM Studio, and MLX load it without custom code. OpenBMB shipped eight artifacts at once: the main model, a b

AMD

Coverage

The vLLM Recipes page documents MiniCPM5-2B as the 2B dense checkpoint in OpenBMB's MiniCPM5 series, built for on-device and resource-constrained deployment with strong performance on agentic tool use, code generation, and reasoning. It uses the standard LlamaForCausalLM architecture, so vLLM loads it natively with no At 2B parameters the BF16 weights (5 GB) fit on a single consumer GPU at TP=1, and the model supports the full native 128K context window, which can be reduced via --max-model-len to free KV cache on smaller GPUs. The page lists supported serving targets including AMD MI300X (192G), MI325X (256G), MI355X (288G), and MI

Videos about MiniCPM5-2B