Sulat.com
AI models
Scaleway logo

Model details

Pixtral 12B 2409

Pixtral 12B is a natively multimodal vision language model combining a 12 billion parameter text decoder with a dedicated 400 million parameter vision encoder. This architecture was trained from the ground up on interleaved image and text data, allowing it to reason across visual and written content in a unified way. The model processes images at their native resolution and aspect ratio, handling variable image sizes and multiple images within a single context without forced standardization. This flexibility makes it well-suited for real-world documents, charts, and mixed-media inputs where visual layouts and token efficiency both matter.

The model delivers leading performance for its weight class across multimodal benchmarks including MMMU (52.5%), ChartQA (81.8%), and DocVQA (90.7%), while maintaining state-of-the-art results on text-only evaluations like MMLU and instruction-following tests. Released under the Apache 2.0 license, it offers open weights for self-hosting and customization. This combination of strong visual reasoning, cost efficiency, and an open license positions Pixtral 12B as a practical choice for applications ranging from document intelligence and data extraction to multimodal conversational systems where both image understanding and text quality are essential.

Scalewaypixtral-12b-2409pixtral

Quick Info

Powered by
Provider
Scaleway
Model key
pixtral-12b-2409
Release date
Sep 25, 2024
Last updated
Mar 17, 2026
Knowledge cutoff
2024-09
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.20

Limits

Output tokens
4,096 tokens
Context window
128,000 tokens

Latest news about Pixtral 12B 2409

No articles yet. Fetch the latest news to show it here.

Videos about Pixtral 12B 2409

More models around Pixtral 12B 2409