Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ambient logo

Model details

Step 3.7 Flash

Step 3.7 Flash is a vision-language model from StepFun designed for multimodal reasoning workloads. It is distributed as an NVIDIA NIM, with an entry under the stepfun-ai namespace on the NVIDIA NGC catalog, which signals a hardware-optimized deployment path for inference. The model is also offered on the Baseten Model Library in a hardware-efficient configuration intended to lower per-token cost and improve autoscaling smoothness for production traffic.

According to Baseten's launch-style announcement, Step 3.7 Flash is a 198-billion-parameter sparse mixture-of-experts vision-language model, a design choice that lets it route tokens to a subset of experts at inference time to balance capability against compute cost. The Baseten post frames the model around multimodal reasoning at scale, suggesting a fit for applications that combine image and text understanding with chain-of-thought style reasoning. For practitioners, the practical story is an open, large-capacity multimodal model that can be served through either an NVIDIA NIM container or a managed Baseten deployment depending on operational preferences.

Ambientstepfun/step-3.7-flash

Quick Info

Powered by
Provider
Ambient
Model key
stepfun/step-3.7-flash
Release date
May 29, 2026
Last updated
May 29, 2026
Knowledge cutoff
2026-03-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.19
Output token cost
$1.14

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Step 3.7 Flash

StepFun (China)

CoverageBenchmark

In third-party benchmark rankings, Step 3.7 Flash holds a composite LLM Stats Score rank of 120, with a top-half placement in coding (99 of 267) and weaker showings in reasoning (117 of 363), math (150 of 327), and tool calling (141 of 194). The model ranks 32nd on Terminal-Bench 2.1 with a score of 0.59, 35th on SWE-B The aggregator notes that benchmark methodology specifics (harness, sampling, pass@k) are not disclosed by the publisher, so these figures should be treated as directional rather than definitive. Reported blended pricing is approximately $0.25 per million tokens, though this reflects marketplace data rather than StepFu

ZenMux

CoverageBenchmark

Step 3.7 Flash achieved top rankings on the Artificial Analysis benchmark in June 2026, securing first place in speed, cost-efficiency, and end-to-end performance while gaining significant traction on OpenRouter and Hugging Face. Output speed reached up to 416 tokens per second—one of the fastest among comparable model Multimodal understanding demonstrations showed the model identifying a dexterous robotic hand from its appearance, recognizing specific joint segments and fingertips, then autonomously searching for and compiling a comprehensive product specification table including manufacturer info, hardware configuration, load capac

Vercel AI Gateway

Coverage

StepFun's Step 3.7 Flash is a sparse mixture-of-experts vision-language model with a 196 billion parameter language backbone and a 1.8 billion parameter vision encoder, activating roughly 11 billion parameters per token. According to coverage citing the Hugging Face model card, the model was released at the end of May Reported throughput reaches up to 400 tokens per second, a figure that still requires independent validation across different hardware and workloads. The active-compute bill, rather than the headline parameter count, is framed as the key practical consideration for developers and startups evaluating local deployment. T

ZenMux

CoverageRelease Notes

StepFun released Step 3.7 Flash, a multimodal Mixture-of-Experts vision-language model targeted at agentic coding and search workflows. According to MarkTechPost's coverage, the model pairs a 196B-parameter language backbone with a separate 1.8B-parameter ViT vision encoder for native image understanding, totaling 198B Step 3.7 Flash offers a 256k-token context window and up to 400 tokens/sec throughput, with the vision encoder injecting image representations into the language backbone rather than running end-to-end fused. The model is positioned for coding agents and search-style agentic use cases that benefit from the sparse-activa

StepFun (China)

Coverage

Step 3.7 Flash is a high-efficiency, production-grade agent model released and open-sourced by StepFun on May 29, 2026 under the Apache 2.0 license. It uses a Sparse Mixture-of-Experts architecture with 196B total language parameters plus a 1.8B vision encoder, activating only about 11B parameters per token, and claims The model supports a 256K context window, native multimodal understanding, internet and visual search enhancement, and tool invocation/orchestration, and offers three selectable inference levels (low, medium, high) with both cloud and on-premises deployment. StepFun reports benchmark scores of 67.1% on ClawEval-1.1 (re

Videos about Step 3.7 Flash