Currently listed through these providers:
Model details
Pixtral Large (latest)
Pixtral Large is a 124-billion parameter multimodal model that brings Mistral's text expertise into the visual domain. Built on the foundation of Mistral Large the listed price, the model pairs a 123-billion parameter multimodal decoder with a 1-billion parameter vision encoder, allowing it to process and reason over documents, charts, natural images, and multilingual text simultaneously. This architecture lets the model tackle advanced image understanding tasks while preserving the strong text capabilities that made its predecessor notable. With the ability to analyze at least 30 high-resolution images within its context window, the model is well-suited for workflows involving dense visual documents or mixed media inputs.
The model emerged as the second entry in Mistral's multimodal series and was released with open weights, available under a research license for non-commercial use and a separate commercial license for business applications. Developers can access it through the Le Chat interface, the API using the pixtral-large-latest endpoint, or download it directly from HuggingFace. Pixtral Large has demonstrated state-of-the-art performance on established multimodal benchmarks including MathVista, DocVQA, and VQAv2, particularly excelling at tasks requiring mathematical reasoning with visual inputs and multilingual optical character recognition. Its design makes it a practical choice for applications that demand both sophisticated text understanding and image comprehension in a single capable model.
Quick Info
Powered by- Provider
- Mistral
- Model key
- pixtral-large-latest
- Release date
- Nov 1, 2024
- Last updated
- Nov 4, 2024
- Knowledge cutoff
- 2024-11
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $2.00
- Output token cost
- $6.00
Limits
- Output tokens
- 128,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare pixtral pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Pixtral Large (latest)
No articles yet. Fetch the latest news to show it here.