Model details
ERNIE 5.0
ERNIE 5.0 is a natively omni-modal foundation model that fundamentally differs from models built on late-fusion pipelines. Rather than stitching together separate vision, audio, and language components, ERNIE 5.0 was designed to jointly model text, images, audio, and video from the earliest stages of training. This architecture enables seamless cross-modal understanding and generation, allowing the model to reason across modalities without the friction of translating between incompatible representations. The model was built to excel in areas where cross-modal reasoning matters: instruction following, factual reasoning, creative writing, agentic planning, and tool use.
The model earned a top-20 position on the competitive LMArena Text Leaderboard, with scores placing it on par with leading frontier models in domains like software engineering and coding. On formal academic benchmarks, ERNIE 5.0 scored 0.87 on the AIME 2025 mathematics evaluation and achieved a rank of 6 on the challenging MMLU-Pro benchmark, which expands multiple-choice options to ten choices and focuses on reasoning-intensive tasks across fourteen domains. These results, validated against Baidu's official scorecard and blog documentation, reflect a model tuned for precision and factual grounding across expert-level question sets in science, law, and healthcare. ERNIE 5.0 represents the latest generation in Baidu's ERNIE series, positioning the model as a foundation for applications requiring deep domain expertise, structured reasoning, and reliable factual recall.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- baidu/ernie-5.0-thinking-preview
- Release date
- Jan 22, 2026
- Last updated
- Jan 22, 2026
- Knowledge cutoff
- 2025-01-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.84
- Output token cost
- $3.37
Limits
- Output tokens
- 64,000 tokens
- Context window
- 128,000 tokens
Latest news about ERNIE 5.0
No articles yet. Fetch the latest news to show it here.