Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Step 3.7 Flash

Step 3.7 Flash is a high-efficiency agent model from StepFun designed for real-world automation, blending native multimodal understanding with reliable tool use. It can interpret images ranging from product interfaces and documents to charts and natural scenes, then write code or invoke tools to act on what it perceives. The release emphasizes "See.Think.Act." workflow, extended web and visual search reach (including long-tail entities and freshly emerged concepts), and stable orchestration across terminals, browsers, and Office tools, with reduced drift on long agent runs. Ecosystem compatibility spans mainstream harnesses such as Claude Code, KiloCode, Hermes Agent, and OpenClaw, which lowers integration friction for teams adopting agentic pipelines.

Positioned around a vendor-claimed throughput of up to 400 TPS, the model targets the latency-sensitive needs of agent loops and is distributed as an open-weight release on GitHub, Hugging Face, and ModelScope under the stepfun-ai namespace, with a packaged NIM available through NVIDIA NGC for streamlined deployment. Its practical fit lies in agentic coding, multi-step research and search workflows, and any application where a vision-capable model must chain tool calls coherently over extended sequences. For teams building production agents that mix vision, search, and tool execution, Step 3.7 Flash offers a compelling balance of open accessibility and agent-first design.

Vercel AI Gatewaystepfun/step-3.7-flash

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
stepfun/step-3.7-flash
Release date
May 29, 2026
Last updated
May 29, 2026
Knowledge cutoff
2026-03-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$1.15

Limits

Input tokens
256,000 tokens
Output tokens
256,000 tokens
Context window
256,000 tokens

Latest news about Step 3.7 Flash

StepFun (China)

CoverageBenchmark

In third-party benchmark rankings, Step 3.7 Flash holds a composite LLM Stats Score rank of 120, with a top-half placement in coding (99 of 267) and weaker showings in reasoning (117 of 363), math (150 of 327), and tool calling (141 of 194). The model ranks 32nd on Terminal-Bench 2.1 with a score of 0.59, 35th on SWE-B The aggregator notes that benchmark methodology specifics (harness, sampling, pass@k) are not disclosed by the publisher, so these figures should be treated as directional rather than definitive. Reported blended pricing is approximately $0.25 per million tokens, though this reflects marketplace data rather than StepFu

ZenMux

CoverageBenchmark

Step 3.7 Flash achieved top rankings on the Artificial Analysis benchmark in June 2026, securing first place in speed, cost-efficiency, and end-to-end performance while gaining significant traction on OpenRouter and Hugging Face. Output speed reached up to 416 tokens per second—one of the fastest among comparable model Multimodal understanding demonstrations showed the model identifying a dexterous robotic hand from its appearance, recognizing specific joint segments and fingertips, then autonomously searching for and compiling a comprehensive product specification table including manufacturer info, hardware configuration, load capac

Vercel AI Gateway

Coverage

StepFun's Step 3.7 Flash is a sparse mixture-of-experts vision-language model with a 196 billion parameter language backbone and a 1.8 billion parameter vision encoder, activating roughly 11 billion parameters per token. According to coverage citing the Hugging Face model card, the model was released at the end of May Reported throughput reaches up to 400 tokens per second, a figure that still requires independent validation across different hardware and workloads. The active-compute bill, rather than the headline parameter count, is framed as the key practical consideration for developers and startups evaluating local deployment. T

ZenMux

CoverageRelease Notes

StepFun released Step 3.7 Flash, a multimodal Mixture-of-Experts vision-language model targeted at agentic coding and search workflows. According to MarkTechPost's coverage, the model pairs a 196B-parameter language backbone with a separate 1.8B-parameter ViT vision encoder for native image understanding, totaling 198B Step 3.7 Flash offers a 256k-token context window and up to 400 tokens/sec throughput, with the vision encoder injecting image representations into the language backbone rather than running end-to-end fused. The model is positioned for coding agents and search-style agentic use cases that benefit from the sparse-activa

StepFun (China)

Coverage

Step 3.7 Flash is a high-efficiency, production-grade agent model released and open-sourced by StepFun on May 29, 2026 under the Apache 2.0 license. It uses a Sparse Mixture-of-Experts architecture with 196B total language parameters plus a 1.8B vision encoder, activating only about 11B parameters per token, and claims The model supports a 256K context window, native multimodal understanding, internet and visual search enhancement, and tool invocation/orchestration, and offers three selectable inference levels (low, medium, high) with both cloud and on-premises deployment. StepFun reports benchmark scores of 67.1% on ClawEval-1.1 (re

Videos about Step 3.7 Flash