Sulat.com
AI models
NEAR AI Cloud logo

Model details

Whisper Large v3

Whisper Large v3 is the third generation of OpenAI's Whisper family, introduced by Alec Radford and colleagues in the paper Robust Speech Recognition via Large-Scale Weak Supervision. It is designed for automatic speech recognition and speech translation, generalising zero-shot across many datasets and domains. The model retains the same Transformer encoder-decoder backbone used in the earlier large and large-v2 checkpoints, with only two small architectural changes: the input spectrogram now uses 128 Mel frequency bins instead of 80, and Cantonese was added as a new language token to broaden multilingual coverage.

Training scale is the main driver of v3's gains: OpenAI combined 1 million hours of weakly labelled audio with 4 million hours of pseudo-labelled audio produced by Whisper large-v2, and trained the model for 2.0 epochs over that mixture, building on the 5 million hours of labelled data lineage behind the family. The result is a 10% to 20% reduction in error rates across a wide variety of languages compared with large-v2, which makes it well suited to multilingual transcription, translation pipelines, and downstream audio analytics. Because the weights are openly published, developers can fine-tune or self-host the model, and it is also offered as a downloadable, DGX Cloud-accelerated NVIDIA NIM for production deployments.

NEAR AI Cloudopenai/whisper-large-v3whisper

Quick Info

Powered by
Provider
NEAR AI Cloud
Model key
openai/whisper-large-v3
Release date
Nov 6, 2023
Last updated
Nov 6, 2023
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
448 tokens
Context window
448 tokens

Latest news about Whisper Large v3

Videos about Whisper Large v3

More models around Whisper Large v3