Currently listed through these providers:
Model details
Voxtral Mini 3B 2507
Voxtral Mini 3B builds directly on Ministral 3B as its language backbone, extending it with dedicated audio understanding capabilities to create a unified model that processes both speech and text without a separate transcription pipeline. The design intent centers on delivering best-in-class automatic speech recognition alongside translation, question answering, and summarization—all from the same compact 3 billion parameter architecture. This approach enables direct function-calling from voice input, letting spoken intents trigger backend workflows or API calls without intermediate steps. The model handles up to 40 minutes of audio for understanding tasks and includes a dedicated transcription mode that prioritizes pure speech-to-text performance, supporting eight of the world's most widely spoken languages with automatic detection.
The training lineage reflects Mistral's strategy of layering multimodal capabilities onto a proven text foundation rather than building audio understanding from scratch. Published under an Apache 2.0 license, the Voxtral family targets both production-scale deployment and local use cases, with the 3B variant designed specifically for compact, efficient operation. The 32K token context window supports long-form audio processing and multi-turn conversations, while built-in Q&A and summarization remove the need for orchestrating separate ASR and language models. For teams that require transcription, translation, or structured analysis of spoken content within a single API call, Voxtral Mini represents a purpose-built option in Mistral's expanding audio model lineup.
Quick Info
Powered by- Provider
- Amazon Bedrock
- Model key
- mistral.voxtral-mini-3b-2507
- Release date
- Dec 1, 2024
- Last updated
- Dec 1, 2024
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.04
- Output token cost
- $0.04
Limits
- Output tokens
- 4,096 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare mistral pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.