Model details
Step 3.7 Flash
Step 3.7 Flash is a 198 billion parameter sparse Mixture-of-Experts vision-language model built around a 196 billion parameter language backbone paired with a 1.8 billion parameter vision encoder, where sparse routing activates only about 11 billion parameters per token. This design is meant to deliver near-frontier reasoning at inference costs closer to a much smaller model, and the release targets coding agents and search-style workflows that benefit from tool use and visual inputs. Independent reviewers describe the model as positioned for agentic tasks, with first-day hands-on testing reporting a 100% tool call success rate in their setup and SWE-Bench PRO scores that outperformed comparable models from other providers.
In qualitative testing, the model has been praised for combining strong agentic reliability with efficient local deployment, including a successful run on compact workstation-class hardware. Reviewers highlighted its value for developer workflows that need visual understanding alongside reasoning, and they noted benchmark results on SWE-Bench PRO and a first-place finish on ClawEval as evidence of practical coding-agent capability. The open-weight availability makes it attractive for teams that want to self-host a capable vision-language model, while the sparse activation pattern keeps per-token compute modest. For practitioners choosing a model, Step 3.7 Flash fits well when the priority is a self-hostable, multimodal reasoning system that can drive tool-using agents and search pipelines without paying full dense-model inference costs.
Quick Info
Powered by- Provider
- Hugging Face
- Model key
- stepfun-ai/Step-3.7-Flash
- Release date
- May 29, 2026
- Last updated
- May 29, 2026
- Knowledge cutoff
- 2026-03-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $1.15
Limits
- Output tokens
- 256,000 tokens
- Context window
- 262,144 tokens