Currently listed through these providers:
Model details
StepAudio 2.5 TTS
StepAudio 2.5 TTS is a proprietary speech generation system positioned as part of the StepAudio family of models designed to move beyond rigid tag-based interfaces toward more intuitive control. Unlike conventional TTS pipelines that rely on markup tags, it accepts plain natural language instructions to steer emotion, pacing, pauses, and overall delivery. The model is also documented as supporting zero-shot voice cloning with full timbre and emotional control, and it operates across Chinese and English, making it useful for multilingual expressive synthesis workflows where natural language cues are preferable to structured tags.
The system's emphasis on three core capabilities—global context control, in-text context control, and zero-sample replication with full-tone control—reflects a design direction aimed at flexible, context-aware speech generation rather than one-off utterance synthesis. Built as a specialized text-to-speech component within StepFun's unified StepAudio 2.5 architecture (described in arXiv:2605.23463), it is offered as a commercial API endpoint, fits teams that want fine-grained expressive control via conversational prompts, and is less suited to strict structured-output or tool-calling scenarios where deterministic formatting is required.
Quick Info
Powered by- Provider
- StepFun (China)
- Model key
- stepaudio-2.5-tts
- Release date
- Apr 16, 2026
- Last updated
- Jul 2, 2026
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens