StepFun
Step 3.5 Flash is now available for public deployment, offering rapid, local large language model inference for developers and organizations.
Model details
Step 3.5 Flash is a foundation model built on a sparse Mixture of Experts architecture, activating only 11B of its 196B total parameters per token. This design targets high inference efficiency, with reported generation speeds of roughly 100–350 tokens per second. StepFun describes it as engineered for frontier reasoning and agentic capabilities while keeping the active compute footprint small, a pattern that places it among the wave of sparse Chinese MoE models emphasizing speed at low active parameter counts.
The model is released under an open weights license and is accompanied by a paper, Hugging Face and ModelScope repositories, GitHub code, and a public chat space, making it practical for builders who want to inspect or deploy it locally. StepFun highlights the model's "intelligence density," claiming its reasoning depth rivals top-tier proprietary systems despite the limited active parameter count. A subsequent open-source drop of the Base checkpoint plus Midtrain artifacts and the SteptronOSS training stack extends the line further, giving downstream teams continuation-pretraining flexibility that is unusual for a release of this scale.
StepFun
Step 3.5 Flash is now available for public deployment, offering rapid, local large language model inference for developers and organizations.
StepFun (China)
StepFun released Step 3.5 Flash on February 5, 2026, as a sparse Mixture-of-Experts model with 196B total parameters and only 11B active parameters, claiming frontier-level reasoning capability while generating at 100–350 tokens per second, according to a ThursdAI release index that links to the StepFun X announcement StepFun followed up on March 5, 2026 by open-sourcing Step 3.5 Flash Base and Midtrain checkpoints, an unusually open release that includes the SteptronOSS training stack on GitHub alongside the weights, giving builders continuation-pretraining flexibility under an Apache-2 oriented license. The same index links to the