Step 3.7 Flash is a sparse Mixture-of-Experts vision-language model that combines a 196B-parameter language backbone with a 1.8B-parameter vision encoder, activating roughly 11B parameters per token for efficiency. This architecture enables native multimodal understanding, letting the model process product UIs, documents, charts, and natural scenes before writing code or calling tools to act on what it sees. The model offers three selectable reasoning levels—low, medium, and high—allowing developers to balance speed, cost, and depth depending on the task. With throughput reaching up to 400 tokens per second and native support for multilingual inputs, the system is built for real-world agents rather than academic benchmarks.
On the SWE-Bench Pro benchmark, which tests software engineering agent performance, Step 3.7 Flash scored 56.3, outperforming the previous Step 3.5 Flash release and competing favorably against larger models like DeepSeek V4 Flash despite activating far fewer parameters per token. The model is available as open-weight GGUF quantizations, ranging from full BF16 precision down to compact formats like Q3 and IQ3, enabling private deployment on workstations with 64–96 GB of unified memory. It integrates with popular agent frameworks including Claude Code, KiloCode, Hermes Agent, OpenClaw, and Skills, reducing the friction of adopting it into existing coding and search workflows. For teams building autonomous coding agents, orchestrating multi-step tool use, or running long-context reasoning pipelines, the combination of strong benchmark performance, efficient MoE design, and broad framework compatibility makes Step 3.7 Flash a practical choice for production agentic systems.