Model details
ERNIE 4.5 VL 28B A3B
The ERNIE 4.5 VL 28B A3B architecture is designed for high efficiency, utilizing a routing mechanism that activates only 3 billion parameters while maintaining the performance levels of larger flagship models. This design allows the model to handle complex tasks, such as visual grounding, video understanding, and STEM problem-solving, with a footprint small enough to run on a single 80 GB GPU. By integrating an innovative Thinking with Images feature, the model can process fine-grained visual details and access long-tail knowledge, making it a versatile tool for developers who require high-performance multimodal reasoning without the overhead of massive parameter counts.
The model underwent an extensive mid-training phase, absorbing a diverse corpus of premium visual-language data to deepen the semantic alignment between modalities. To achieve its reasoning capabilities, the training process employed multimodal reinforcement learning on verifiable tasks, utilizing GSPO and IcePop strategies to stabilize its mixture-of-experts training alongside dynamic difficulty sampling. These post-training advancements, combined with improved instruction-following, provide a robust foundation for commercial applications. Its open-source license and compatibility with standard inference frameworks like vLLM and Transformers ensure it remains a practical, forward-looking choice for developers building sophisticated visual-textual applications.
Quick Info
Powered by- Provider
- NovitaAI
- Model key
- baidu/ernie-4.5-vl-28b-a3b
- Release date
- Jun 30, 2025
- Last updated
- Jun 14, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.56
Limits
- Output tokens
- 8,000 tokens
- Context window
- 30,000 tokens