Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Step 3.7 Flash

The model overview is temporarily unavailable.

DevPass (LLM Gateway)step-3.7-flash

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
step-3.7-flash
Release date
May 29, 2026
Last updated
May 29, 2026
Knowledge cutoff
2026-03-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$1.15

Limits

Output tokens
256,000 tokens
Context window
262,144 tokens

Latest news about Step 3.7 Flash

DevPass (LLM Gateway)

Coverage

Startup Fortune reports that StepFun's Step 3.7 Flash, released at the end of May 2026 under Apache 2.0 by the Shanghai-based AI lab (backed by investors including Tencent), is drawing attention for pairing open weights with strong claimed benchmark performance and lower active compute. The 196B-parameter language back According to StepFun's Hugging Face model card, the model supports a 256k context window, three reasoning levels, and throughput of up to 400 tokens per second, claims that explain rapid local-inference community interest. StepFun is positioning Step 3.7 Flash as infrastructure for agents and production deployments rat

DevPass (LLM Gateway)

CoverageRelease Notes

MarkTechPost reports that StepFun released Step 3.7 Flash on May 29, 2026, as a multimodal 198B-parameter sparse Mixture-of-Experts vision-language model targeting agentic use cases, with native vision input and improved tool-use reliability over Step 3.5 Flash. The architecture pairs a 196B language backbone with a se Key specifications include a 256k context window, throughput of up to 400 tokens/sec, three reasoning levels (low, medium, high), and Apache 2.0 licensing. The model is designed for production-level Agents, with design priorities balancing speed, cost, reliable execution, and complex task-handling capabilities rather t

DevPass (LLM Gateway)

CoverageBenchmark

BenchLM's aggregator page (data as of September 23, 2026) lists Step 3.7 Flash with a capability score of 41.3/100 (field median 50.4), ranking 109th of 196 tracked models. Reported metrics include a $0.20 input / $1.15 output price per million tokens, a blended price of $0.67, a speed of 196 tok/s (field median 91), a Category rankings include Agentic at rank 56 of 105 (percentile 47th, score 35.1 across 7 verified benchmarks, weighted 22%) and Coding at rank 78 of 135 (percentile 43rd, score 35.9 across 2 verified benchmarks, weighted 20%), with Multimodal at score 71.0 across 2 verified benchmarks. The page notes that Reasoning, K

DevPass (LLM Gateway)

Coverage

Lambda's deployment page details hardware throughput benchmarks for Step 3.7 Flash using vLLM with Multi-Token Prediction speculative decoding, measured at 8192 in / 2048 out tokens and 32 concurrent requests. Performance figures include 7,982 tok/s total throughput on 4× NVIDIA B200 GPUs (TTFT 2,183 ms, ITL 52.7 ms), The page confirms Step 3.7 Flash as StepFun AI's 198-billion-parameter sparse MoE vision-language model released open-weight under Apache 2.0, activating ~11B parameters per token with a 256K context and three reasoning modes. The architecture pairs a 196B language backbone (carried over nearly unchanged from Step 3.5

DevPass (LLM Gateway)

Coverage

Step 3.7 Flash is a high-efficiency AI model released and open-sourced by StepFun on May 29, 2026, designed for production-level Agents and positioned as a successor within StepFun's Flash series (predecessor: Step 3.5 Flash). It employs a Sparse Mixture-of-Experts architecture with 196B language parameters plus a 1.8B Key capabilities include native multimodal understanding and execution, internet and visual search enhancement, tool invocation and orchestration, a 256k context length, and three selectable inference levels (low, medium, high). The model supports both cloud and on-premises deployment, and scored 67.1% on ClawEval-1.1

Videos about Step 3.7 Flash