Sulat.com
AI models
Vercel AI Gateway logo

Model details

Pixtral 12B

Pixtral 12B represents a significant step into multimodal processing, designed to handle both natural images and complex documents with high precision. At its core, the model utilizes a custom-built vision encoder developed from scratch, which allows it to ingest visual data at its natural resolution and aspect ratio. This architecture provides users with greater flexibility regarding token usage when processing images, ensuring that the model remains highly capable in multimodal tasks without sacrificing its foundational strength as a high-performance text model.

The model stands out for its ability to maintain competitive performance against much larger alternatives, often surpassing significantly bigger models in various benchmarks. By integrating advanced vision capabilities with a robust language foundation, it excels at interpreting multiple images within a single session. This design lineage emphasizes efficiency and versatility, making it a powerful tool for developers looking for a high-performing, compact solution that does not compromise on quality when navigating complex visual and textual inputs.

Vercel AI Gatewaymistral/pixtral-12bpixtral

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
mistral/pixtral-12b
Release date
Sep 1, 2024
Last updated
Sep 1, 2024
Knowledge cutoff
2024-09
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.15

Limits

Output tokens
128,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare pixtral pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Pixtral 12B

Videos about Pixtral 12B

More models around Pixtral 12B