Currently listed through these providers:
Model details
Pixtral 12B
Pixtral 12B represents a significant step into multimodal processing, designed to handle both natural images and complex documents with high precision. At its core, the model utilizes a custom-built vision encoder developed from scratch, which allows it to ingest visual data at its natural resolution and aspect ratio. This architecture provides users with greater flexibility regarding token usage when processing images, ensuring that the model remains highly capable in multimodal tasks without sacrificing its foundational strength as a high-performance text model.
The model stands out for its ability to maintain competitive performance against much larger alternatives, often surpassing significantly bigger models in various benchmarks. By integrating advanced vision capabilities with a robust language foundation, it excels at interpreting multiple images within a single session. This design lineage emphasizes efficiency and versatility, making it a powerful tool for developers looking for a high-performing, compact solution that does not compromise on quality when navigating complex visual and textual inputs.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- mistral/pixtral-12b
- Release date
- Sep 1, 2024
- Last updated
- Sep 1, 2024
- Knowledge cutoff
- 2024-09
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.15
Limits
- Output tokens
- 128,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare pixtral pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.