Model details
Step 3.7 Flash
Step 3.7 Flash is a vision-language model built around a sparse Mixture-of-Experts design, pairing a roughly 196B parameter language backbone with a separate vision encoder while activating only about 11B parameters per token. That routing lets the model handle long multimodal inputs, including images and video alongside text, while keeping per-request inference costs closer to those of much smaller systems. It ships as an open-weight release, which means teams that want full control can run and fine-tune it locally rather than relying solely on a hosted endpoint, and it exposes tool calling plus configurable reasoning controls for embedding it into agentic pipelines.
In hands-on coding-agent testing the model stood out for tool reliability and software engineering benchmarks, reportedly achieving a perfect tool-call success rate and a SWE-Bench PRO score that surpassed competing flash-tier models, with a leading position on the ClawEval agent benchmark. Those results, combined with vision and video understanding in the same checkpoint, make it a practical fit for agent workflows that need to look at screenshots, diagrams, or recorded screen context, then plan multi-step actions against external tools. The combination of a sparse activation pattern and an open-weights license makes it especially appealing for organizations that want frontier-leaning reasoning quality without absorbing the cost of fully dense inference at similar scale.
Quick Info
Powered by- Provider
- StepFun (China)
- Model key
- step-3.7-flash
- Release date
- May 29, 2026
- Last updated
- Jun 29, 2026
- Knowledge cutoff
- 2026-03-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.185
- Output token cost
- $1.11
Limits
- Input tokens
- 256,000 tokens
- Output tokens
- 256,000 tokens
- Context window
- 256,000 tokens