Currently listed through these providers:
Model details
StepAudio 2.5 ASR
StepAudio 2.5 ASR is positioned within StepFun's audio model family as a dedicated speech-to-text endpoint, designed to convert spoken audio into accurate written transcripts without requiring developers to split files into smaller segments. According to third-party documentation, the model is invoked through a single API request that can handle recordings of roughly five to thirty minutes, using an approximately 32K context window in one pass rather than stitching together multiple shorter transcriptions. This single-call design removes the client-side chunking logic that older transcription pipelines required, letting applications send a full meeting, interview, or podcast segment and receive a coherent transcript back from one request.
A separate "StepAudio 2.5 Technical Report" has been indexed by a third-party literature review service, signaling that StepFun has published a more detailed account of the broader StepAudio 2.5 line, though the supplied excerpt does not expose architecture, training data, or benchmark figures from that report. In practice, the ASR variant fits workflows that need efficient long-form transcription across mixed Chinese and English content, offering developers a streamlined path from raw audio to usable text. Teams building transcription features for podcasts, lectures, call recordings, or multilingual media archives can benefit from the simplified single-request approach, while anyone needing deeper claims about parameter counts, training scale, or evaluation results will need to consult the underlying technical report directly.
Quick Info
Powered by- Provider
- StepFun (Global)
- Model key
- stepaudio-2.5-asr
- Release date
- Apr 24, 2026
- Last updated
- Jul 2, 2026
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens