Currently listed through these providers:
Model details
Voxtral Small 24B
Voxtral Small 24B is part of Mistral's Voxtral family of multimodal audio chat models, trained to understand both spoken audio and text documents while preserving strong text capabilities. The model is released under the Apache 2.0 license and is small enough to run locally, making it a practical option for developers who want audio understanding without relying on closed cloud services. Its design pairs a 24-billion-parameter backbone with a 32K token context window, allowing the model to process audio clips up to roughly 40 minutes long and sustain extended multi-turn conversations that mix voice and text inputs.
In benchmark evaluations presented in the Voxtral paper, Voxtral Small outperforms a number of closed-source alternatives across a diverse range of audio tasks while maintaining competitive text performance. The release also contributed three new benchmarks aimed at measuring speech understanding on knowledge and trivia, helping the community compare audio-aware language models more fairly. For practitioners, this translates into a flexible open-weight model that suits speech-to-text transcription pipelines, voice-enabled assistants, and multimodal applications where audio comprehension, long context handling, and the freedom to self-host matter more than raw scale.
Quick Info
Powered by- Provider
- GreenPT
- Model key
- voxtral-small-24b-2507
- Release date
- Jul 15, 2025
- Last updated
- Jul 15, 2025
- Knowledge cutoff
- 2025-07
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.228
- Output token cost
- $0.513
Limits
- Output tokens
- 16,384 tokens
- Context window
- 32,768 tokens