Sulat.com
AI models
Scaleway logo

Model details

Whisper Large v3

Whisper Large v3 carries forward the architecture established in its predecessors, with two targeted updates: the spectrogram frontend expanded to 128 Mel frequency bins from 80, and a dedicated language token added for Cantonese. The model was designed as a general-purpose automatic speech recognition and translation system that could handle a broad range of audio conditions and languages without task-specific fine-tuning. This architectural stability across versions reflects a focus on reliable, zero-shot generalization rather than constant redesign.

The large-v3 checkpoint was trained on a dataset assembled from one million hours of weakly labeled audio supplemented by four million hours of pseudo-labeled data generated using the v2 model, with the full mixture processed for two complete epochs. This distillation-like approach yielded measurable gains, cutting word error rates by 10 to 20 percent compared to v2 across a wide variety of languages. The resulting model functions as an open, deployable foundation for transcription, voicebot pipelines, podcast processing, and accessibility tooling, where its multilingual coverage and strong out-of-the-box accuracy reduce the need for domain-specific adaptation.

Scalewaywhisper-large-v3whisper

Quick Info

Powered by
Provider
Scaleway
Model key
whisper-large-v3
Release date
Sep 1, 2023
Last updated
Mar 17, 2026
Knowledge cutoff
2023-09
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
0 tokens

Latest news about Whisper Large v3

No articles yet. Fetch the latest news to show it here.

Videos about Whisper Large v3

More models around Whisper Large v3