Currently listed through these providers:
Model details
Pixtral Large (latest)
Built on Mistral Large the listed price as a 124-billion parameter multimodal decoder paired with a one-billion parameter vision encoder, Pixtral Large is the second entry in Mistral's multimodal series and succeeds the earlier Pixtral 12B. Its design preserves the pure text reasoning of Mistral Large the listed price while adding native image understanding, so the same model handles long-form language work and visual analysis. Reported results place it at the top of standard multimodal benchmarks such as MathVista, DocVQA, and VQAv2, signaling a clear lineage from Mistral's text-first foundation work toward a unified vision-language model.
The the cataloged API limit context window lets Pixtral Large absorb roughly thirty images in a single prompt, which makes it well suited to workflows that mix dense text with diagrams, scanned documents, charts, or photos, such as research review, financial reporting, or technical documentation. Weights are released openly under the Mistral Research License for research and educational use, with a separate commercial license available for production deployments. Practical fit centers on tasks where both reading comprehension and visual interpretation matter, especially when long contexts and heterogeneous inputs need to be reasoned over together.
Quick Info
Powered by- Provider
- Merge Gateway
- Model key
- mistral/pixtral-large-latest
- Release date
- Nov 1, 2024
- Last updated
- Nov 4, 2024
- Knowledge cutoff
- 2024-11
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $2.00
- Output token cost
- $6.00
Limits
- Output tokens
- 128,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare pixtral pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.