Step 3.5 Flash is StepFun's open-source foundation model built on a sparse Mixture-of-Experts architecture, which selectively activates only 11 billion of its 196 billion total parameters per token. This approach prioritizes what StepFun describes as "intelligence density" over brute parameter scale, allowing the model to deliver reasoning depth while keeping per-token computational cost manageable. Published weights are hosted on Hugging Face under the stepfun-ai organization, giving researchers and developers direct access to the model. The design emphasis on activating a small expert subset per token is central to its performance profile and distinguishes it from dense transformer counterparts.
Positioned as a reasoning model optimized for agent and code workflows, Step 3.5 Flash targets scenarios that demand both depth of inference and sustained throughput at long context lengths. StepFun's Step Plan platform exposes the model via dedicated API paths, and documentation describes it as high-speed inference tuned for intelligent agent applications. A derivative checkpoint, step-3.5-flash-2603, builds on this foundation with further optimizations for high-frequency agent scenarios, offering improved token efficiency, faster inference, and an optional low-reasoning mode for cost-sensitive deployments. This variant family suggests a roadmap focused on practical agent deployment rather than purely academic benchmark leadership.