evroc
Voxtral Small 24B 2507 pricing: $0.10/M input, $0.30/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.
Model details
Voxtral Small 24B is positioned as a multimodal extension of the Mistral Small 3 text model, layering state-of-the-art audio understanding on top of best-in-class text performance so a single model can handle both written and spoken inputs. Rather than being a separate architecture from scratch, it inherits the text competency of its Mistral Small lineage and adds native audio input capability, making it well suited to workflows that combine voice and document data. Practical strengths highlighted in third-party descriptions include speech transcription, speech translation, and general audio understanding, which means the same weights can serve transcription pipelines, multilingual voice assistants, and assistants that need to reason over recorded meetings or calls alongside text context.
Because the weights are open, the model can be self-hosted for latency-sensitive or privacy-conscious audio workloads, and it is also distributed through hosted inference channels, with an official model card published on AWS Bedrock that documents its availability as a managed offering. The audio-input, text-output design lets teams route voice data into existing text-oriented pipelines and prompts, which is a natural fit for retrieval, summarization, and analytics over conversational recordings. Compared with pure-text Llms, the differentiator is the integrated audio encoder; compared with pure ASR systems, the differentiator is the retained ability to reason about the transcript in context, so it is best matched to use cases where both transcription quality and downstream language understanding matter.
evroc
Voxtral Small 24B 2507 pricing: $0.10/M input, $0.30/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.