Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Mistral logo

Model details

Voxtral Small

Voxtral Small builds on Mistral's Small 3 text foundation by layering in audio input capabilities, making it a multimodal model designed to handle speech alongside traditional language tasks. According to the model's description, it incorporates state-of-the-art audio input handling while retaining the text performance of its predecessor, positioning it as a unified model for teams that need both language understanding and speech processing without stitching together separate systems. The open-weights release under the mistralai/Voxtral-Small-24B-2507 Hugging Face repository makes the full model accessible for self-hosting, fine-tuning, and research use, giving developers the freedom to adapt it to specialized audio-text workflows.

In practice, Voxtral Small is aimed at workloads that revolve around audio: speech transcription, translation across languages, and broader audio understanding such as interpreting spoken content in context. Its text-generation capabilities remain central, so it can transcribe or translate audio and then summarize, answer questions about, or act on that content within a single model call. The combination of open weights, audio-native input, and continued text strengths makes it a practical fit for organizations building voice-enabled assistants, multilingual media tools, meeting and podcast pipelines, and accessibility applications that need a single capable model to bridge spoken and written language.

Mistralvoxtral-small-2507voxtral

Quick Info

Powered by
Provider
Mistral
Model key
voxtral-small-2507
Release date
Jul 15, 2025
Last updated
Jul 15, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
32,000 tokens
Context window
32,768 tokens

Latest news about Voxtral Small

Mistral

Official sourceAnnouncement

Mistral's official blog post announced the Voxtral family on July 15, 2025, explicitly naming Voxtral Small as the 24B parameter variant for production-scale applications alongside a 3B Voxtral Mini for local and edge deployments. Both models ship under Apache 2.0 and are also served via Mistral's API, with a transcribe-optimized Voxtral Mini Transcribe variant routing transcription queries for cost and latency efficiency. The post highlights technical capabilities including a 32k token context handling up to 30 minutes of audio for transcription or 40 minutes for understanding, built-in Q&A and summarization, and automatic multilingual detection. Mistral positions Voxtral Small as matching Scribe-class transcription quality from closed APIs while costing less than half the price, targeting enterprise speech-intelligence deployments.

Mistral

Official sourceDocumentation

Mistral's official documentation page introduces Voxtral Small v25.07 (voxtral-small-2507) as a generally available audio-input instruct model released July 15, 2025 under the Apache 2.0 license. The page confirms a 32k token context window and lists API pricing of $0.004 per minute for audio, $0.1 per million input tokens, and $0.4 per million output tokens, giving developers clear cost guidance for production deployment. The same documentation lists supported endpoints and developer features including /v1/chat/completions, /v1/conversations, and /v1/batch, alongside structured outputs, function calling, document Q&A, prefix support, and batching. These capabilities position Voxtral Small 2507 as a production-ready speech-understanding model with flexible integration paths for Mistral's chat and conversation APIs.

Videos about Voxtral Small

More models around Voxtral Small