Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

StepFun 3.5 Flash

StepFun 3.5 Flash is positioned as a sparse Mixture-of-Experts reasoning model that activates only a fraction of its total capacity per token, with roughly 11 billion of 196 billion parameters engaged during inference. This selective routing is the central architectural idea, allowing the model to behave like a much larger system on paper while keeping the compute footprint closer to a mid-sized model at runtime. The same design choice underpins its emphasis on reasoning and tool-oriented workflows, making it well suited to structured problem solving, retrieval-heavy pipelines, and bilingual Chinese-English workloads where long, careful reasoning chains are common.

In practical terms, the model offers a very large 262,114-token context and output window, which makes it attractive for document analysis, code repositories, and other long-form tasks that would strain smaller context limits. Third-party reviewers describe it as delivering a substantial share of frontier model quality at a fraction of the inference cost, particularly when routed through a hybrid setup that mixes it with a more expensive model for harder prompts. Teams comfortable with self-hosting can also explore GGUF quantizations for additional savings, while those preferring managed access get implicit caching, tool use, and streaming integration through the gateway's standard SDK patterns.

Vercel AI Gatewaystepfun/step-3.5-flash

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
stepfun/step-3.5-flash
Release date
Jan 29, 2026
Last updated
Feb 13, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.09
Output token cost
$0.30

Limits

Output tokens
262,114 tokens
Context window
262,114 tokens

Latest news about StepFun 3.5 Flash

No articles yet. Fetch the latest news to show it here.

Videos about StepFun 3.5 Flash