Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

Step 3.7 Flash

Step 3.7 Flash from StepFun is a vision-language model built around a sparse Mixture-of-Experts design, with roughly 198 billion total parameters in the language backbone and an additional compact vision encoder, yet only about 11 billion parameters activate per token. This routing strategy aims to keep inference costs closer to a small model while still supporting a very long the cataloged API limit token context window, three selectable reasoning settings, and native handling of text, image, and video inputs. The model is also packaged as an NVIDIA NIM container under the stepfun-ai organization on the NGC catalog, which makes it straightforward to drop into existing GPU-based serving pipelines and local workstations such as the DGX Spark for private deployments.

In hands-on agentic testing, reviewers reported a perfect tool-call success rate over a full day of workflows and competitive scores on software engineering benchmarks, including 56.3 on SWE-Bench PRO—said to outperform comparable flash-tier models—and a claimed first-place 67.1 on ClawEval. Because the weights are available, teams can self-host for cost control, fine-tune for proprietary code or document pipelines, or route traffic through ZenMux for managed access. The combination of sparse activation, long context, multimodal input, and reliable tool execution makes Step 3.7 Flash a practical fit for production agents, codebase assistants, and document-heavy retrieval workflows where frontier-class reasoning at low active-parameter cost matters more than raw peak model size.

ZenMuxstepfun/step-3.7-flash

Quick Info

Powered by
Provider
ZenMux
Model key
stepfun/step-3.7-flash
Release date
May 29, 2026
Last updated
May 29, 2026
Knowledge cutoff
2026-03-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$1.15

Limits

Input tokens
256,000 tokens
Output tokens
256,000 tokens
Context window
256,000 tokens

Latest news about Step 3.7 Flash

StepFun (China)

CoverageBenchmark

In third-party benchmark rankings, Step 3.7 Flash holds a composite LLM Stats Score rank of 120, with a top-half placement in coding (99 of 267) and weaker showings in reasoning (117 of 363), math (150 of 327), and tool calling (141 of 194). The model ranks 32nd on Terminal-Bench 2.1 with a score of 0.59, 35th on SWE-B The aggregator notes that benchmark methodology specifics (harness, sampling, pass@k) are not disclosed by the publisher, so these figures should be treated as directional rather than definitive. Reported blended pricing is approximately $0.25 per million tokens, though this reflects marketplace data rather than StepFu

ZenMux

CoverageBenchmark

Step 3.7 Flash achieved top rankings on the Artificial Analysis benchmark in June 2026, securing first place in speed, cost-efficiency, and end-to-end performance while gaining significant traction on OpenRouter and Hugging Face. Output speed reached up to 416 tokens per second—one of the fastest among comparable model Multimodal understanding demonstrations showed the model identifying a dexterous robotic hand from its appearance, recognizing specific joint segments and fingertips, then autonomously searching for and compiling a comprehensive product specification table including manufacturer info, hardware configuration, load capac

Vercel AI Gateway

Coverage

StepFun's Step 3.7 Flash is a sparse mixture-of-experts vision-language model with a 196 billion parameter language backbone and a 1.8 billion parameter vision encoder, activating roughly 11 billion parameters per token. According to coverage citing the Hugging Face model card, the model was released at the end of May Reported throughput reaches up to 400 tokens per second, a figure that still requires independent validation across different hardware and workloads. The active-compute bill, rather than the headline parameter count, is framed as the key practical consideration for developers and startups evaluating local deployment. T

ZenMux

CoverageRelease Notes

StepFun released Step 3.7 Flash, a multimodal Mixture-of-Experts vision-language model targeted at agentic coding and search workflows. According to MarkTechPost's coverage, the model pairs a 196B-parameter language backbone with a separate 1.8B-parameter ViT vision encoder for native image understanding, totaling 198B Step 3.7 Flash offers a 256k-token context window and up to 400 tokens/sec throughput, with the vision encoder injecting image representations into the language backbone rather than running end-to-end fused. The model is positioned for coding agents and search-style agentic use cases that benefit from the sparse-activa

StepFun (China)

Coverage

Step 3.7 Flash is a high-efficiency, production-grade agent model released and open-sourced by StepFun on May 29, 2026 under the Apache 2.0 license. It uses a Sparse Mixture-of-Experts architecture with 196B total language parameters plus a 1.8B vision encoder, activating only about 11B parameters per token, and claims The model supports a 256K context window, native multimodal understanding, internet and visual search enhancement, and tool invocation/orchestration, and offers three selectable inference levels (low, medium, high) with both cloud and on-premises deployment. StepFun reports benchmark scores of 67.1% on ClawEval-1.1 (re

Videos about Step 3.7 Flash