Model details
StepFun 3.5 Flash
StepFun 3.5 Flash is positioned as a sparse Mixture-of-Experts reasoning model that activates only a fraction of its total capacity per token, with roughly 11 billion of 196 billion parameters engaged during inference. This selective routing is the central architectural idea, allowing the model to behave like a much larger system on paper while keeping the compute footprint closer to a mid-sized model at runtime. The same design choice underpins its emphasis on reasoning and tool-oriented workflows, making it well suited to structured problem solving, retrieval-heavy pipelines, and bilingual Chinese-English workloads where long, careful reasoning chains are common.
In practical terms, the model offers a very large 262,114-token context and output window, which makes it attractive for document analysis, code repositories, and other long-form tasks that would strain smaller context limits. Third-party reviewers describe it as delivering a substantial share of frontier model quality at a fraction of the inference cost, particularly when routed through a hybrid setup that mixes it with a more expensive model for harder prompts. Teams comfortable with self-hosting can also explore GGUF quantizations for additional savings, while those preferring managed access get implicit caching, tool use, and streaming integration through the gateway's standard SDK patterns.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- stepfun/step-3.5-flash
- Release date
- Jan 29, 2026
- Last updated
- Feb 13, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.09
- Output token cost
- $0.30
Limits
- Output tokens
- 262,114 tokens
- Context window
- 262,114 tokens
Latest news about StepFun 3.5 Flash
No articles yet. Fetch the latest news to show it here.