Step 3.5 Flash is a foundation model built on a sparse Mixture of Experts architecture, activating only 11B of its 196B total parameters per token. This design targets high inference efficiency, with reported generation speeds of roughly 100–350 tokens per second. StepFun describes it as engineered for frontier reasoning and agentic capabilities while keeping the active compute footprint small, a pattern that places it among the wave of sparse Chinese MoE models emphasizing speed at low active parameter counts.
The model is released under an open weights license and is accompanied by a paper, Hugging Face and ModelScope repositories, GitHub code, and a public chat space, making it practical for builders who want to inspect or deploy it locally. StepFun highlights the model's "intelligence density," claiming its reasoning depth rivals top-tier proprietary systems despite the limited active parameter count. A subsequent open-source drop of the Base checkpoint plus Midtrain artifacts and the SteptronOSS training stack extends the line further, giving downstream teams continuation-pretraining flexibility that is unusual for a release of this scale.