Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
EmpirioLabs AI logo

Model details

Step 3.7 Flash

Step 3.7 Flash is positioned as a vision-language model released by StepFun and surfaced on the NVIDIA NGC catalog as a NIM-deployable artifact under the stepfun-ai team, giving teams a path to run it on optimized infrastructure. A MarkTechPost announcement frames it around coding agents and search workflows, suggesting an emphasis on tool-augmented reasoning, retrieval-style tasks, and developer-oriented applications where combining visual inputs with code and language understanding matters.

The MarkTechPost headline describes Step 3.7 Flash as a 198B Mixture-of-Experts vision-language model, pointing to a large but sparsely activated design that can balance capacity with inference efficiency. Its listing in the NVIDIA NGC catalog as a NIM entry indicates vendor packaging for streamlined deployment, which is useful for organizations that want a vision-capable model integrated into agentic or search-oriented pipelines without building custom serving stacks from scratch.

EmpirioLabs AIstep-3-7-flash

Quick Info

Powered by
Provider
EmpirioLabs AI
Model key
step-3-7-flash
Release date
May 29, 2026
Last updated
May 29, 2026
Knowledge cutoff
2026-03-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$1.15

Limits

Input tokens
256,000 tokens
Output tokens
131,072 tokens
Context window
256,000 tokens

Latest news about Step 3.7 Flash

EmpirioLabs AI

CoverageBenchmark

Step 3.7 Flash appears on the LLM Stats leaderboard ranked 127th overall, with capability-tier standing of average in general, below the top half in tool calling, legal, and finance categories, and coding/reasoning/math in the middle tier. Its blended price point is reported at approximately $0.25 per million tokens wi Benchmark evaluations sourced from StepFun's official scorecard include Terminal-Bench 2.1 (rank 34, score 0.59/1), SWE-Bench Pro (rank 36, score 0.56/1), and Humanity's Last Exam with tools in text-only mode (rank 8, score 0.47/1). The methodology notes indicate these are StepFun-internal or publisher-reported agentic

EmpirioLabs AI

Coverage

StepFun released Step 3.7 Flash at the end of May 2026 under an Apache 2.0 license, positioning it as an open-weights, inference-efficient alternative in the emerging "Flash" category of AI models. The model is a sparse mixture-of-experts vision-language system with a 196 billion parameter language backbone paired with According to StepFun's Hugging Face model card, Step 3.7 Flash supports a 256k context window, offers three selectable reasoning levels, and is claimed to deliver throughput of up to 400 tokens per second. The release is framed for practical local deployment and agent workflows, with the active-parameter figure highlig

EmpirioLabs AI

CoverageRelease Notes

StepFun released Step 3.7 Flash on May 29, 2026, as a 198B-parameter sparse Mixture-of-Experts vision-language model targeting agentic use cases. It pairs a 196B-parameter language backbone with a 1.8B-parameter vision encoder (ViT) for native image understanding, activating approximately 11B parameters per token durin Key specifications include a 256k-token context window, claimed throughput of up to 400 tokens per second, three reasoning levels (low, medium, high), and an Apache 2.0 license. The vision encoder runs as a separate 1.8B ViT module that injects image representations into the language backbone, enabling multimodal agent

EmpirioLabs AI

Coverage

Step 3.7 Flash was released and open-sourced by StepFun on May 29, 2026, as a high-efficiency large model in the Flash series succeeding Step 3.5 Flash. It targets production-level agent deployment, balancing speed, cost, reliable execution, and complex task handling as agent technology transitions from demos to enterp Technical specifications confirm a Sparse MoE architecture with 196B language parameters plus a 1.8B ViT vision encoder, activating only 11B parameters per token and reaching up to 400 tokens/s. It supports a 256k context length and three inference levels (low, medium, high), with both cloud and on-premises deployment

Videos about Step 3.7 Flash