Currently listed through these providers:
Model details
Pixtral 12B 2409
Pixtral 12B is a natively multimodal vision language model combining a 12 billion parameter text decoder with a dedicated 400 million parameter vision encoder. This architecture was trained from the ground up on interleaved image and text data, allowing it to reason across visual and written content in a unified way. The model processes images at their native resolution and aspect ratio, handling variable image sizes and multiple images within a single context without forced standardization. This flexibility makes it well-suited for real-world documents, charts, and mixed-media inputs where visual layouts and token efficiency both matter.
The model delivers leading performance for its weight class across multimodal benchmarks including MMMU (52.5%), ChartQA (81.8%), and DocVQA (90.7%), while maintaining state-of-the-art results on text-only evaluations like MMLU and instruction-following tests. Released under the Apache 2.0 license, it offers open weights for self-hosting and customization. This combination of strong visual reasoning, cost efficiency, and an open license positions Pixtral 12B as a practical choice for applications ranging from document intelligence and data extraction to multimodal conversational systems where both image understanding and text quality are essential.
Quick Info
Powered by- Provider
- Scaleway
- Model key
- pixtral-12b-2409
- Release date
- Sep 25, 2024
- Last updated
- Mar 17, 2026
- Knowledge cutoff
- 2024-09
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.20
Limits
- Output tokens
- 4,096 tokens
- Context window
- 128,000 tokens
Latest news about Pixtral 12B 2409
No articles yet. Fetch the latest news to show it here.