Sulat.com
AI models
StepFun (China) logo

Model details

Step 3.7 Flash

Step 3.7 Flash is a vision-language model built around a sparse Mixture-of-Experts design, pairing a roughly 196B parameter language backbone with a separate vision encoder while activating only about 11B parameters per token. That routing lets the model handle long multimodal inputs, including images and video alongside text, while keeping per-request inference costs closer to those of much smaller systems. It ships as an open-weight release, which means teams that want full control can run and fine-tune it locally rather than relying solely on a hosted endpoint, and it exposes tool calling plus configurable reasoning controls for embedding it into agentic pipelines.

In hands-on coding-agent testing the model stood out for tool reliability and software engineering benchmarks, reportedly achieving a perfect tool-call success rate and a SWE-Bench PRO score that surpassed competing flash-tier models, with a leading position on the ClawEval agent benchmark. Those results, combined with vision and video understanding in the same checkpoint, make it a practical fit for agent workflows that need to look at screenshots, diagrams, or recorded screen context, then plan multi-step actions against external tools. The combination of a sparse activation pattern and an open-weights license makes it especially appealing for organizations that want frontier-leaning reasoning quality without absorbing the cost of fully dense inference at similar scale.

StepFun (China)step-3.7-flash

Quick Info

Powered by
Provider
StepFun (China)
Model key
step-3.7-flash
Release date
May 29, 2026
Last updated
Jun 29, 2026
Knowledge cutoff
2026-03-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.185
Output token cost
$1.11

Limits

Input tokens
256,000 tokens
Output tokens
256,000 tokens
Context window
256,000 tokens

Latest news about Step 3.7 Flash

Videos about Step 3.7 Flash

Recent tweets and retweets from StepFun (China)