SiliconFlow
Step 3.5 Flash, the most capable StepFun AI's open-source foundation model is here. It is engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency.
Model details
Step-3.5-Flash is a foundation model engineered to provide deep reasoning and agentic capabilities while maintaining high operational efficiency. It utilizes a sparse Mixture of Experts architecture, which contains 196 billion total parameters but selectively activates only 11 billion per token. This design intent focuses on intelligence density, allowing the model to handle complex mathematical logic and code-related tasks with a level of performance that rivals top-tier proprietary systems. By balancing depth with a lightweight active parameter count, it serves as a robust brain for autonomous agents that require both accuracy and speed.
To achieve its rapid response times, the model incorporates MTP-3 technology, which enables the prediction of three tokens simultaneously to double inference efficiency. It also employs a sliding window with global attention to manage extensive text sequences, ensuring stability during long-chain logical tasks. These architectural choices, combined with hybrid attention schedules, allow the model to maintain high throughput in production environments. Its design is specifically optimized for real-time interactions, making it a practical choice for developers building complex, long-context applications that demand both cost-effective compute and advanced reasoning.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
SiliconFlow
Step 3.5 Flash, the most capable StepFun AI's open-source foundation model is here. It is engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency.
SiliconFlow
An independent newsletter report dated 2026-02-03 covers StepFun's release of Step 3.5 Flash, describing it as a sparse mixture-of-experts model with 196B total parameters and roughly 11B active per token, released under an Apache 2.0 open-weights licence. The article cites StepFun-reported benchmark figures of 97.3% o The report positions the model as prioritizing strong reasoning within an efficient context window and fast inference rather than raw parameter count, situating the release within a competitive Lunar New Year open-model wave from Chinese labs. For Step-3.5-Flash specifically, the piece is useful third-party documentati
SiliconFlow
StepFun's organization page on Hugging Face lists a dedicated Step-3.5-Flash collection of Text Generation models, with entries at 199B parameters dated to February and March 2026 and community engagement of roughly 106k likes on the top entry, providing creator-attributed confirmation that stepfun-ai/Step-3.5-Flash is Although the excerpt is a directory view rather than a model card or technical report, it surfaces direct artifacts developers can use: Text Generation checkpoints in the 197B–199B parameter range, an Apache-2.0 open-weights lineage under the stepfun-ai namespace, and links to interactive Streamlit chatbot demos hosted
SiliconFlow
Compare GLM-4.7 and Step-3.5-Flash across performance, cost, capabilities, and real-world use cases. See which model fits your needs.
SiliconFlow
Step-3.5-Flash: Open-source MoE model for advanced reasoning, coding, and agentic tasks. RL-powered, high performance.