Step 3.7 Flash is StepFun's response to the demand for fast, agent-friendly multimodal models that can also think step by step. It is built as a sparse Mixture-of-Experts vision-language system with roughly 198 billion total parameters, of which only about 11 billion activate per token, paired with a 1.8 billion parameter vision encoder on top of a 196 billion parameter language backbone. This routing design lets the model deliver reasoning depth on par with much larger dense models while keeping inference cost closer to a small model. Its open weights are published as stepfun-ai/Step-3.7-Flash on Hugging Face, and it accepts text, images, and video alongside standard temperature controls, tool calling, and structured reasoning modes.
In practical use, Step 3.7 Flash is positioned for agentic coding, tool-driven workflows, and multimodal prompts where speed matters as much as accuracy. Independent hands-on testing on an NVIDIA DGX Spark reported a 100% tool-call success rate, a SWE-Bench PRO score of 56.3, and a ClawEval score of 67.1 that placed it ahead of competing flash-tier peers from other labs. With a long the cataloged API limit context window and a free ZenMux routing option, it is a strong fit for developers who want to run vision-aware agents, code assistants, or research pipelines without managing paid inference or running the model themselves, while still benefiting from frontier-style reasoning.