Sulat.com
AI models
Privatemode AI logo

Model details

Whisper large-v3

Whisper Large-v3 is OpenAI's open-source automatic speech recognition system, designed to convert spoken audio into text and to translate non-English speech into English. It is built as a large transformer-based encoder-decoder trained on a broad and diverse corpus of multilingual and multitask supervised data, which is what enables a single checkpoint to handle both transcription and translation across many languages. With 1,550 million parameters, it represents the largest configuration in the Whisper family and is positioned for demanding, noise-heavy conditions where robustness matters as much as raw accuracy.

In practice, Whisper Large-v3 is well suited for production transcription pipelines that need multilingual coverage, accepting common audio formats and producing timestamped output at both word and segment granularity for downstream indexing, subtitling, or analytics. Its open weights make it straightforward to self-host or fine-tune for domain-specific vocabulary, accents, or audio conditions, and the model's roughly 10 percent word error rate has kept it competitive enough that newer systems, such as Microsoft's MAI-Transcribe-1, are explicitly benchmarked against it rather than simply replacing it. Developers choosing this variant should expect the strongest general-purpose Whisper accuracy available, balanced against higher compute requirements than the smaller Whisper checkpoints.

Privatemode AIwhisper-large-v3whisper

Quick Info

Powered by
Provider
Privatemode AI
Model key
whisper-large-v3
Release date
Oct 1, 2024
Last updated
Oct 1, 2024
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
4,096 tokens
Context window
448 tokens

Latest news about Whisper large-v3

Privatemode AI

CoverageRelease Notes

Microsoft's MAI-Transcribe-1 model currently leads the FLEURS benchmark in 11 core languages and outperforms competitors like OpenAI's Whisper-large-v3 and Google's Gemini 3.1 Flash., Microsoft's MAI-Transcribe-1 model currently leads the FLEURS benchmark in 11 core languages and outperforms competitors like OpenAI's W

Privatemode AI

Official sourceRelease Notes

Privatemode's release notes provide direct operational context for developers using Whisper large-v3 via `/v1/audio/transcriptions`: v1.45.0 introduced a 1-hour decoded-audio limit per transcription request and added rate limits for cached tokens and audio file size, and pre-v1.46.0 versions are now retired as of Augus The same notes flag an API behavior change affecting OpenAI-compatible transcription responses: starting with v1.54.0, the `duration` field in verbose JSON transcription responses is encoded as a string, and v1.55.0 will switch it to a JSON number to match the OpenAI API. For Whisper large-v3 (and Voxtral Mini 3B) call

Videos about Whisper large-v3

More models around Whisper large-v3