Currently listed through these providers:
Model details
Voxtral Small 24B 2507
Voxtral Small is a multimodal speech-and-text model from Mistral AI that brings audio understanding to a compact, locally runnable footprint. Built as an enhancement of Mistral Small 3, it adds state-of-the-art audio input capabilities while retaining strong text performance. The model can operate in a dedicated transcription mode, supports voice-driven function calling, and includes built-in Q&A and summarization without needing separate speech-to-text and language model pipelines. It natively handles multilingual audio, automatically detecting language and transcribing or translating across a broad set of spoken languages.
The Voxtral family was trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a range of audio benchmarks while outperforming several closed-source models. Voxtral Small handles long-form audio content suitable for extended interviews, meetings, or podcasts, and is released with open weights under the Apache 2.0 license. This makes it a practical choice for teams that want to run a capable speech model on their own infrastructure while still having access to managed cloud inference through platforms like Amazon Bedrock.
Quick Info
Powered by- Provider
- Amazon Bedrock
- Model key
- mistral.voxtral-small-24b-2507
- Release date
- Jul 1, 2025
- Last updated
- Jul 1, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.35
Limits
- Output tokens
- 8,192 tokens
- Context window
- 32,000 tokens
Transparent token rates
Compare mistral pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Voxtral Small 24B 2507
No articles yet. Fetch the latest news to show it here.