Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

ERNIE-4.5-VL-28B-A3B-Thinking

This model utilizes a Mixture-of-Experts architecture, combining a 21-billion parameter language decoder with a 7-billion parameter vision encoder to process complex visual and textual data. By activating only 3 billion parameters per token, it achieves a balance between high-performance reasoning and computational efficiency. The design is specifically intended to handle nuanced visual-textual tasks, such as extracting on-screen text, analyzing video content over time, and executing tool-based actions like image zooming or searching to resolve fine-grained details.

The model underwent an extensive mid-training phase, absorbing a diverse corpus of visual-language reasoning data to deepen semantic alignment between modalities. It incorporates multimodal reinforcement learning, utilizing GSPO and IcePop strategies to stabilize its routing architecture alongside dynamic difficulty sampling for improved learning efficiency. These post-training advancements enable the model to excel in STEM problem-solving and visual grounding, providing a robust, Apache 2.0-licensed foundation for developers seeking to integrate sophisticated multimodal analysis into commercial applications.

NovitaAIbaidu/ernie-4.5-vl-28b-a3b-thinking

Quick Info

Powered by
Provider
NovitaAI
Model key
baidu/ernie-4.5-vl-28b-a3b-thinking
Release date
Nov 26, 2025
Last updated
Nov 26, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.39
Output token cost
$0.39

Limits

Output tokens
65,536 tokens
Context window
131,072 tokens

Latest news about ERNIE-4.5-VL-28B-A3B-Thinking

NovitaAI

Coverage

Baidu open-sourced ERNIE-4.5-VL-28B-A3B-Thinking on November 13, 2025, releasing it as an Apache 2.0 visual reasoning model. The architecture is a Mixture-of-Experts design with 28 billion total parameters but only 3 billion active during inference, making it computationally efficient while targeting vision-language ta The release made weights available on Hugging Face with an accompanying GitHub repository and Baidu blog post, reinforcing Baidu's open-weights multimodal strategy. The model is part of a broader wave of Chinese open model releases, alongside siblings like ERNIE 4.5 and later ERNIE-Image and ERNIE 5.1 Preview. This ind

Videos about ERNIE-4.5-VL-28B-A3B-Thinking