Model details
Step 3.7 Flash
Step 3.7 Flash is StepFun's high-efficiency multimodal release, structured as a sparse Mixture-of-Experts model pairing a roughly 196-billion-parameter language backbone with a vision encoder that activates only about 11 billion parameters per token. That sparse routing is the core design idea: deliver near-frontier language and vision reasoning while keeping per-token compute closer to a small model. Native image and video understanding are built in rather than bolted on, and the model exposes selectable reasoning effort levels so callers can trade depth against latency and cost depending on the task at hand.
In practical terms, Step 3.7 Flash is positioned for coding agents, structured-output pipelines, and long-context productivity work, with tool calling as a first-class workflow. A third-party review on an NVIDIA DGX Spark reported a perfect tool-call success rate in agentic testing, along with SWE-Bench PRO and ClawEval scores that placed it ahead of comparable Flash-tier rivals. Weights are open and hosted under the stepfun-ai organization on Hugging Face, making it a strong fit for teams that want to self-host a multimodal MoE for agent stacks, document or video analysis, or extended-context assistants without paying frontier-model inference prices.
Quick Info
Powered by- Provider
- UnoRouter
- Model key
- step-3.7-flash:free
- Release date
- May 29, 2026
- Last updated
- May 29, 2026
- Knowledge cutoff
- 2026-03-01
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Input tokens
- 256,000 tokens
- Output tokens
- 256,000 tokens
- Context window
- 256,000 tokens