Model details
Step 3.7 Flash
Step 3.7 Flash is a frontier multimodal reasoning model aimed squarely at agentic workloads, pairing native image and video understanding with text output that can write code or invoke tools. On Step Plan, it is exposed as the flagship reasoning entry and supports three configurable reasoning effort levels (low, medium, and high) so teams can trade depth against latency and token spend. The launch emphasizes real-world acting rather than passive description: the model reads product UIs, documents, charts, and natural-scene images, then drives terminals, browsers, Office tools, and web or visual search to finish the task, with an advertised ceiling of up to 400 TPS for tight agent loops.
Weights are openly published across GitHub, Hugging Face, and ModelScope under the StepFun organization, and community reports show the model running on local llama.cpp setups such as an NVIDIA DGX Spark, which keeps integration costs low for self-hosted pipelines. The release is positioned as a step up from Step 3.5 Flash, posting a SWE-Bench Pro score of 56.3 at 196B parameters against 51.3 for its predecessor and edging out DeepSeek V4 Flash's 55.6 in the same comparison, illustrating stronger agentic coding headroom at the same parameter scale. Compatibility with mainstream harnesses including Claude Code, KiloCode, Hermes Agent, OpenClaw, and Skills means existing orchestration stacks can adopt the model without retooling, making it a practical fit for teams building long-running search, browser, or coding agents that need reliable tool calling.
Quick Info
Powered by- Provider
- StepFun Step Plan (Global)
- Model key
- step-3.7-flash
- Release date
- May 29, 2026
- Last updated
- May 29, 2026
- Knowledge cutoff
- 2026-03-01
- Input modalities
- Output modalities
- Capabilities
Limits
- Input tokens
- 256,000 tokens
- Output tokens
- 256,000 tokens
- Context window
- 256,000 tokens