Step 3.7 Flash is a sparse Mixture-of-Experts vision-language model positioned as a "Flash-tier" efficiency release rather than a maximum-intelligence frontier run. StepFun describes it as a high-efficiency model aimed at real-world agents, with the headline marketing claim of up to 400 tokens per second inference throughput, and a parameter count of roughly 196B parameters carried forward from the prior Flash generation. The model is multimodal at its core, accepting images across the full range of product UIs, documents, charts, and natural scenes, and is explicitly designed to translate visual understanding into action by writing code or invoking tools.
In practice, Step 3.7 Flash is tuned for agentic workflows rather than standalone chat. StepFun highlights reliable tool use and orchestration across terminals, browsers, Office tools, and search backends, with reduced drift and fewer broken tool calls on long-running sessions, plus compatibility with mainstream agent harnesses including Claude Code, KiloCode, Hermes Agent, and OpenClaw. Internal agentic-coding benchmarks reported on the launch page place the model ahead of the previous Step 3.5 Flash on SWE-Bench Pro, framing the release as a measurable step up in coding-agent competence for a Flash-class budget. The combination of a 262K-token context window, low input pricing, multimodal grounding, and agent-harness support makes it a natural fit for teams building production assistants that need to read screens, call tools, and iterate quickly without paying frontier-tier rates.