Sulat.com
AI models
GreenPT logo

Model details

Voxtral Small 24B

Voxtral Small 24B is part of Mistral's Voxtral family of multimodal audio chat models, trained to understand both spoken audio and text documents while preserving strong text capabilities. The model is released under the Apache 2.0 license and is small enough to run locally, making it a practical option for developers who want audio understanding without relying on closed cloud services. Its design pairs a 24-billion-parameter backbone with a 32K token context window, allowing the model to process audio clips up to roughly 40 minutes long and sustain extended multi-turn conversations that mix voice and text inputs.

In benchmark evaluations presented in the Voxtral paper, Voxtral Small outperforms a number of closed-source alternatives across a diverse range of audio tasks while maintaining competitive text performance. The release also contributed three new benchmarks aimed at measuring speech understanding on knowledge and trivia, helping the community compare audio-aware language models more fairly. For practitioners, this translates into a flexible open-weight model that suits speech-to-text transcription pipelines, voice-enabled assistants, and multimodal applications where audio comprehension, long context handling, and the freedom to self-host matter more than raw scale.

GreenPTvoxtral-small-24b-2507mistral

Quick Info

Powered by
Provider
GreenPT
Model key
voxtral-small-24b-2507
Release date
Jul 15, 2025
Last updated
Jul 15, 2025
Knowledge cutoff
2025-07
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.228
Output token cost
$0.513

Limits

Output tokens
16,384 tokens
Context window
32,768 tokens

Latest news about Voxtral Small 24B

Videos about Voxtral Small 24B

More models around Voxtral Small 24B