Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

Step 3.7 Flash

Step 3.7 Flash is a sparse Mixture-of-Experts vision-language model that combines a 196B-parameter language backbone with a 1.8B-parameter vision encoder, activating roughly 11B parameters per token for efficiency. This architecture enables native multimodal understanding, letting the model process product UIs, documents, charts, and natural scenes before writing code or calling tools to act on what it sees. The model offers three selectable reasoning levels—low, medium, and high—allowing developers to balance speed, cost, and depth depending on the task. With throughput reaching up to 400 tokens per second and native support for multilingual inputs, the system is built for real-world agents rather than academic benchmarks.

On the SWE-Bench Pro benchmark, which tests software engineering agent performance, Step 3.7 Flash scored 56.3, outperforming the previous Step 3.5 Flash release and competing favorably against larger models like DeepSeek V4 Flash despite activating far fewer parameters per token. The model is available as open-weight GGUF quantizations, ranging from full BF16 precision down to compact formats like Q3 and IQ3, enabling private deployment on workstations with 64–96 GB of unified memory. It integrates with popular agent frameworks including Claude Code, KiloCode, Hermes Agent, OpenClaw, and Skills, reducing the friction of adopting it into existing coding and search workflows. For teams building autonomous coding agents, orchestrating multi-step tool use, or running long-context reasoning pipelines, the combination of strong benchmark performance, efficient MoE design, and broad framework compatibility makes Step 3.7 Flash a practical choice for production agentic systems.

Nvidiastepfun-ai/step-3.7-flashdeprecated

Quick Info

Powered by
Provider
Nvidia
Model key
stepfun-ai/step-3.7-flash
Release date
May 28, 2026
Last updated
May 28, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
16,384 tokens
Context window
256,000 tokens

Latest news about Step 3.7 Flash

Nvidia

Coverage

StepFun's Step 3.7 Flash release at the end of May 2026 under an Apache 2.0 license is drawing attention because it combines open weights, multimodal capability, and low active compute in a package developers can actually try locally. The Shanghai-based lab, backed by investors including Tencent, released the model to According to StepFun's Hugging Face model card as cited in coverage, Step 3.7 Flash supports a 256K context window, three reasoning levels, and throughput of up to 400 tokens per second. The sparse MoE vision-language system has a 196B-parameter language backbone and a 1.8B-parameter vision encoder, activating about 11

Nvidia

CoverageRelease Notes

StepFun released Step 3.7 Flash on May 29, 2026, a 198-billion-parameter sparse Mixture-of-Experts vision-language model targeting agentic use cases, adding native vision input and improved tool-use reliability over Step 3.5 Flash. It pairs a 196B-parameter language backbone with a 1.8B-parameter vision encoder (ViT) f Key published specifications include a 256K-token context window, up to 400 tokens/sec throughput, and three configurable reasoning levels (low, medium, high). StepFun positions Step 3.7 Flash as infrastructure for agentic coding and search workflows rather than as a raw scale leader. The MoE architecture keeps inferen

Hugging Face

Coverage

Lambda's deployment documentation explicitly names Step 3.7 Flash, attributes it to StepFun AI, and confirms the Apache 2.0 release with a 198B-parameter vision-language architecture, 11B active parameters per token, 256K context, and low/medium/high reasoning modes. The language backbone carries over from Step 3.5 Fla Lambda's own vLLM-MTP benchmarks show aggregate throughput of 7,982 tok/s on 4× B200, 6,224 tok/s on 8× H100, and 2,370 tok/s on 8× A100 under an 8192-in/2048-out, 32-concurrent workload. Architecture specifics include 288 routed plus 1 shared expert with top-8 sigmoid routing, 45 layers, 4096 hidden size, and S3F1 hyb

Hugging Face

Coverage

Step 3.7 Flash is described as a high-efficiency large AI model aimed at production-level Agents, open-sourced by StepFun on May 29, 2026 under the Apache 2.0 license, with Step 3.5 Flash identified as its predecessor in the Flash series. It uses a sparse MoE architecture totaling 196B+1.8B (ViT) parameters and activat The model delivers native multimodal understanding and execution plus internet and visual search enhancement, along with tool invocation and orchestration, and reportedly scored 67.1% on ClawEval-1.1 and 49.5% on Toolathlon for multi-tool collaboration. Following release, domestic hardware vendors including Tianshu Zhi

Videos about Step 3.7 Flash