Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

GPT-4o Mini Transcribe

GPT-4o Mini Transcribe is a speech-to-text model from OpenAI released in March 2025 as part of a new generation of audio models that also includes the larger GPT-4o Transcribe and the customizable GPT-4o Mini TTS. VentureBeat's coverage describes it as one of three proprietary voice models unveiled together, built to let developers add speech recognition to existing text-based applications through the OpenAI API. Microsoft Azure's developer blog corroborates the model identity, noting the same three-model lineup becoming available in public preview on Azure AI Foundry, with GPT-4o Mini Transcribe positioned as the more compact sibling to GPT-4o Transcribe for transcription workloads.

As a lighter member of OpenAI's transcription family, GPT-4o Mini Transcribe is aimed at developers who need reliable audio-to-text conversion without stepping up to the full GPT-4o Transcribe model. The Azure source frames it as a speech-to-text model that outperforms previous benchmarks in its class, making it well-suited for voice input features in chat interfaces, meeting transcription, call analytics, and other applications where spoken content needs to flow into text-based pipelines. The VentureBeat report reinforces this practical orientation by noting the model is available through the API for third-party developers building their own apps, alongside a custom demo site for limited testing. Together these sources paint the model as a focused, developer-friendly transcription tool rather than a general-purpose language model.

DevPass (LLM Gateway)gpt-4o-mini-transcribegpt

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
gpt-4o-mini-transcribe
Release date
Mar 20, 2025
Last updated
Mar 20, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.25
Output token cost
$5.00

Limits

Output tokens
16,000 tokens
Context window
16,000 tokens

Transparent token rates

Compare GPT-4o Mini Transcribe pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT-4o Mini Transcribe

DevPass (LLM Gateway)

CoverageBenchmark

The VoiceCodeBench benchmark, designed by besimple.ai to evaluate how well automatic speech recognition models transcribe critical workflow entities (URLs, commands, file paths, postal addresses, email addresses, dates, numbers, etc.), explicitly lists GPT 4o Mini Transcribe (Streaming) among the 18 evaluated systems. The benchmark notes that STT models still struggle with separator- and punctuation-heavy values — URLs (46% mean recovery), commands (51%), file paths (52%) — while handling dates, plain numbers, percentages, measurements, and phone extensions at 98-99%. The top performer was GPT Live Transcribe at 67.67% TSR, with GPT

DevPass (LLM Gateway)

CoverageAnalysis

Artificial Analysis provides benchmarking and analysis of OpenAI GPT-4o family speech-to-text models across API providers, using its AA-WER v2 Word Error Rate Index measured across three datasets (AA-AgentTalk 50%, VoxPopuli-Cleaned-AA 25%, Earnings22-Cleaned-AA 25%). The methodology note explicitly names GPT-4o Mini T The page aggregates WER, speed factor, and price comparisons across 35 of 59 models and multiple API providers (OpenAI, Deepgram, Google, Microsoft Azure, AWS Bedrock, ElevenLabs, and others). Pricing is quoted in USD per 1000 minutes of audio, with the analysis plotting accuracy versus cost and speed trade-offs for ea

Videos about GPT-4o Mini Transcribe

More models around GPT-4o Mini Transcribe