Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

Step 3.5 Flash

Step 3.5 Flash is built on a sparse Mixture-of-Experts architecture that sets it apart through what its developers call "intelligence density"—the ability to deliver deep, frontier-level reasoning while remaining computationally lean. The model activates only 11 billion of its 196 billion parameters per token, allowing it to rival the reasoning depth of top-tier proprietary systems without the corresponding computational overhead. This design philosophy prioritizes sharp, reliable reasoning alongside fast execution, making it particularly well-suited for agentic workflows where the model must plan, adapt, and take action across extended contexts.

The model has been evaluated against rigorous benchmarks that underscore its practical strengths, achieving strong scores on mathematical reasoning and software engineering task resolution that place it among the most capable open-weight alternatives available. As an Apache 2.0 licensed release, Step 3.5 Flash is openly accessible to developers who want to inspect, fine-tune, or deploy it across different infrastructure—from cloud endpoints to local workstations capable of running quantized versions. Its combination of open availability, competitive benchmark performance, and efficiency-focused architecture positions it as a practical choice for teams building autonomous agents, coding assistants, or any application requiring sophisticated reasoning under real-time constraints.

Nvidiastepfun-ai/step-3.5-flashdeprecated

Quick Info

Powered by
Provider
Nvidia
Model key
stepfun-ai/step-3.5-flash
Release date
Feb 2, 2026
Last updated
Feb 2, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
16,384 tokens
Context window
256,000 tokens

Latest news about Step 3.5 Flash

StepFun (Global)

Coverage

StepFun released Step 3.5 Flash on February 5, 2026 as a 196B-parameter sparse Mixture-of-Experts model with only 11B active parameters, claiming frontier-level reasoning performance while generating at 100–350 tokens per second. The release continued a trend of sparse Chinese MoE models that deliver high throughput at On March 5, 2026, StepFun followed up by open-sourcing Step 3.5 Flash Base along with Midtrain checkpoints and the SteptronOSS training stack, an unusually open release that ships training artifacts alongside the weights. The panel highlighted its Apache-2 orientation and called the continuation-pretraining flexibility

Videos about Step 3.5 Flash