Step 5 Preview is positioned by StepFun as the flagship entry in its reasoning lineup, with the official documentation describing it as built for agentic work and exposed through a dedicated reasoning integration on the Step Plan API. It accepts text, image, and video inputs and operates within a very large context window that lets long-running agents and multimodal pipelines keep substantial working memory in a single session. StepFun groups it alongside the step-3.7-flash reasoning model in the same API surface, suggesting a shared effort-level control scheme that lets developers tune how much deliberation the model invests per turn, a useful knob for trading latency against quality in production agent loops.
The model sits within StepFun's broader sparse-MoE family tradition, where the step-3.5-flash sibling uses a 196B-total / 11B-activated mixture-of-experts design that the lab has emphasized for high-speed agent and coding tasks, and Step 5 Preview's flagship agentic framing implies similar efficiency-oriented goals at a higher capability tier. On external evaluation trackers, it draws attention for particularly strong multimodal and grounded workflows involving screenshots, documents, and charts, making it a practical fit for teams that need an agent that can read complex visual inputs and reason about them. Developers building tool-using assistants, document-understanding pipelines, or long-horizon research agents get a multimodal front end, reasoning-effort controls, and a context window broad enough to hold sizable corpora, all served through one consistent Step Plan endpoint.