Currently listed through these providers:
Model details
Voxtral Small 24B 2507
Voxtral Small 24B 2507 is Mistral's multimodal audio chat model that combines spoken-audio comprehension with strong text handling, positioned as an audio-enabled evolution of the Mistral Small 3 lineage. The accompanying technical report describes it as trained to understand both speech and text documents, with claims of state-of-the-art results across a diverse range of audio benchmarks while preserving the text competence of its predecessor. The model is presented as compact enough to run locally, making it attractive for teams that want multimodal capabilities without depending on a large hosted system, and both Voxtral variants are released under the Apache 2.0 license, which fits the model's positioning as a practical, open alternative to closed-source audio chat systems.
Beyond audio, the model retains the practical conveniences expected of a modern chat system, including tool calling and structured outputs, which makes it suitable for building agents and assistants that need to convert speech or audio input into machine-readable actions. The 32K context window is large enough to handle audio files of roughly forty minutes and to sustain long multi-turn conversations, broadening its usefulness for transcription, meeting summarization, voice-driven workflows, and extended Q&A sessions. For practitioners, the model is a reasonable fit when an application needs both robust text generation and direct audio understanding in a single, locally deployable package, rather than stitching together separate speech-to-text and language model services.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- mistralai/voxtral-small-24b-2507
- Release date
- Oct 30, 2025
- Last updated
- Oct 30, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.30
Limits
- Output tokens
- 26,214 tokens
- Context window
- 32,768 tokens