DevPass (LLM Gateway)
Startup Fortune reports that StepFun's Step 3.7 Flash, released at the end of May 2026 under Apache 2.0 by the Shanghai-based AI lab (backed by investors including Tencent), is drawing attention for pairing open weights with strong claimed benchmark performance and lower active compute. The 196B-parameter language back According to StepFun's Hugging Face model card, the model supports a 256k context window, three reasoning levels, and throughput of up to 400 tokens per second, claims that explain rapid local-inference community interest. StepFun is positioning Step 3.7 Flash as infrastructure for agents and production deployments rat