Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
StepFun Step Plan (China) logo

Model details

Step 3.7 Flash

Step 3.7 Flash is designed as a high-efficiency model for real-world agents, combining visual understanding with action-oriented tool use. It can interpret product interfaces, documents, charts, natural scenes, and video, then generate code or invoke tools based on what it sees. Its workflow focus includes web and visual search, orchestration across terminals, browsers, office tools, and sustained multi-step agent runs.

The model is positioned for agentic coding, multimodal reasoning, and integration with mainstream agent harnesses, including Claude Code, KiloCode, Hermes Agent, OpenClaw, and Skills. The official materials emphasize reliable orchestration, less drift, fewer broken tool calls, and fewer failed runs during longer tasks. The cited SWE-Bench Pro result is 56.3, compared with 51.3 for Step 3.5 Flash and 55.6 for DeepSeek V4 Flash in the same excerpt, supporting its fit for coding workflows and tool-driven applications.

StepFun Step Plan (China)step-3.7-flash

Quick Info

Powered by
Provider
StepFun Step Plan (China)
Model key
step-3.7-flash
Release date
May 29, 2026
Last updated
May 29, 2026
Knowledge cutoff
2026-03-01
Input modalities
Output modalities
Capabilities

Limits

Input tokens
256,000 tokens
Output tokens
256,000 tokens
Context window
256,000 tokens

Latest news about Step 3.7 Flash

StepFun Step Plan (China)

Coverage

Startup Fortune reports that Shanghai-based StepFun, backed by investors including Tencent, released Step 3.7 Flash at the end of May 2026 under an Apache 2.0 license as a sparse mixture-of-experts vision-language system. According to the article citing StepFun's Hugging Face model card, the model has a 196 billion par The piece also states that Step 3.7 Flash supports a 256K context window, three reasoning levels, and claimed throughput of up to 400 tokens per second, positioning the release as infrastructure for agent workflows and local deployment scenarios. It frames StepFun as a serious efficiency-focused competitor in the Chine

StepFun Step Plan (China)

CoverageRelease Notes

MarkTechPost reports that on May 29, 2026, StepFun released Step 3.7 Flash, a 198B-parameter sparse Mixture-of-Experts vision-language model designed for coding agents and search workflows. The model pairs a 196B-parameter language backbone with a 1.8B-parameter vision encoder (ViT), adding native vision input and impr The article frames Step 3.7 Flash as StepFun's entry into efficiency-focused, multimodal MoE models aimed at developers building agentic pipelines. It emphasizes coding-agent reliability and search-workflow integration as concrete target workloads, while positioning the release as a step up from the Step 3.5 Flash gene

StepFun Step Plan (China)

Coverage

The Baidu Baike entry documents Step 3.7 Flash as a StepFun model released and open-sourced on May 29, 2026 under the Apache 2.0 license, following its predecessor Step 3.5 Flash. It specifies a Sparse MoE architecture with 196B language parameters plus a 1.8B ViT vision encoder, activating roughly 11B parameters per t Reported benchmark scores include 67.1% on ClawEval-1.1 (daily autonomous task execution) and 49.5% on Toolathlon (multi-tool collaboration), with capabilities spanning native multimodal understanding, internet and visual search enhancement, and tool invocation/orchestration. The entry also notes that companies includi

Videos about Step 3.7 Flash