Step 3.7 Flash from StepFun is a vision-language model built around a sparse Mixture-of-Experts design, with roughly 198 billion total parameters in the language backbone and an additional compact vision encoder, yet only about 11 billion parameters activate per token. This routing strategy aims to keep inference costs closer to a small model while still supporting a very long the cataloged API limit token context window, three selectable reasoning settings, and native handling of text, image, and video inputs. The model is also packaged as an NVIDIA NIM container under the stepfun-ai organization on the NGC catalog, which makes it straightforward to drop into existing GPU-based serving pipelines and local workstations such as the DGX Spark for private deployments.
In hands-on agentic testing, reviewers reported a perfect tool-call success rate over a full day of workflows and competitive scores on software engineering benchmarks, including 56.3 on SWE-Bench PRO—said to outperform comparable flash-tier models—and a claimed first-place 67.1 on ClawEval. Because the weights are available, teams can self-host for cost control, fine-tune for proprietary code or document pipelines, or route traffic through ZenMux for managed access. The combination of sparse activation, long context, multimodal input, and reliable tool execution makes Step 3.7 Flash a practical fit for production agents, codebase assistants, and document-heavy retrieval workflows where frontier-class reasoning at low active-parameter cost matters more than raw peak model size.