Mistral
The first multi-modal, text+image-to-text model from Mistral AI. Its weights were launched via torrent: https://x.com/mistralai/status/1833758285167722836.
Model details
Pixtral 12B is a 12-billion-parameter multimodal model engineered to bridge the gap between visual perception and natural language understanding. At its core, the model features a vision encoder built from scratch, which enables it to ingest images at their native resolutions and aspect ratios. This architectural choice provides users with significant flexibility regarding token usage when processing visual data. Unlike many models that prioritize one modality over the other, this design ensures that the system maintains high-level text performance while simultaneously excelling at interpreting complex documents and images.
The model demonstrates a strong lineage of performance, consistently outperforming other open-source models of similar size, such as Llama-3.2 11B and Qwen-2-VL 7B. Its capability to handle any number of images within its extensive context window makes it a practical tool for tasks requiring deep visual analysis. By delivering results that rival much larger models—even those seven times its size—it serves as a highly efficient solution for developers seeking to integrate advanced multimodal reasoning into their applications without sacrificing the linguistic quality expected from a cutting-edge text model.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Mistral
The first multi-modal, text+image-to-text model from Mistral AI. Its weights were launched via torrent: https://x.com/mistralai/status/1833758285167722836.
Mistral
Pixtral 12B pricing: $0.10/M input, $0.10/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.
This exact model name is also listed by 3 other providers.