Model details
Step-3
Step-3 is a large multimodal reasoning model built on a sparse Mixture-of-Experts architecture with 321B total parameters and 38B active parameters per token. This design allows the model to selectively engage only a fraction of its capacity for each computation, dramatically reducing decoding costs while maintaining strong performance across vision-language tasks. The architecture incorporates Multi-Matrix Factorization Attention and Attention-FFN Disaggregation techniques to further optimize efficiency across diverse hardware, from flagship accelerators to lower-end devices. The model's design emphasis on grounded multimodal reasoning aims to improve visual interpretation accuracy and reduce hallucination issues common in vision-language systems.
Step-3 positions itself as a cost-effective alternative to proprietary systems while delivering competitive benchmark results across multiple reasoning and knowledge tasks. Available as open-source with weights on Hugging Face and GitHub, the model offers transparency for developers and researchers seeking to deploy multimodal AI without reliance on closed APIs. The architecture's efficiency advantages make it practical for production deployments where both performance and computational economics matter. By open-sourcing the model weights, StepFun enables the research community to study, fine-tune, and extend the model across diverse applications, from document understanding to visual question answering.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- stepfun/step-3
- Release date
- Jul 31, 2025
- Last updated
- Jul 31, 2025
- Knowledge cutoff
- 2025-01-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.21
- Output token cost
- $0.57
Limits
- Output tokens
- 64,000 tokens
- Context window
- 65,536 tokens
Latest news about Step-3
No articles yet. Fetch the latest news to show it here.