Sulat.com
AI models
StepFun (China) logo

Model details

StepAudio 2.5 TTS

StepAudio 2.5 TTS is a proprietary speech generation system positioned as part of the StepAudio family of models designed to move beyond rigid tag-based interfaces toward more intuitive control. Unlike conventional TTS pipelines that rely on markup tags, it accepts plain natural language instructions to steer emotion, pacing, pauses, and overall delivery. The model is also documented as supporting zero-shot voice cloning with full timbre and emotional control, and it operates across Chinese and English, making it useful for multilingual expressive synthesis workflows where natural language cues are preferable to structured tags.

The system's emphasis on three core capabilities—global context control, in-text context control, and zero-sample replication with full-tone control—reflects a design direction aimed at flexible, context-aware speech generation rather than one-off utterance synthesis. Built as a specialized text-to-speech component within StepFun's unified StepAudio 2.5 architecture (described in arXiv:2605.23463), it is offered as a commercial API endpoint, fits teams that want fine-grained expressive control via conversational prompts, and is less suited to strict structured-output or tool-calling scenarios where deterministic formatting is required.

StepFun (China)stepaudio-2.5-ttsstep

Quick Info

Powered by
Provider
StepFun (China)
Model key
stepaudio-2.5-tts
Release date
Apr 16, 2026
Last updated
Jul 2, 2026
Input modalities
Output modalities
Capabilities
Base catalog fields only

Limits

Output tokens
0 tokens
Context window
0 tokens

Latest news about StepAudio 2.5 TTS

Videos about StepAudio 2.5 TTS

Recent tweets and retweets from StepFun (China)

More models around StepAudio 2.5 TTS