LongCat-2.5 Preview is Meituan's multimodal follow-up in the LongCat family, built as a preview release focused on long-range agent tasks rather than a general chat upgrade. It preserves the mixture-of-experts footprint established in LongCat-2.0, carrying forward roughly 1.6 trillion total parameters with about 48 billion activated per pass, and keeps the million-token context window that LongCat-2.0 used for sweeping over documents, code repositories, logs, and multi-step workflows. The headline change is native multimodality: the model can read images directly, handling cross-modal question answering, executive summarization of visual content, and more involved visual reasoning, which lets a single model reason over both screen captures and structured text artifacts in long agent runs.
The release is positioned squarely at long-process automation in agent-style environments such as terminals, browsers, desktop software, spreadsheets, and design tools, with coding emphasized as a leading use case. Meituan ships the preview alongside OpenAI- and Anthropic-compatible API endpoints, compatibility paths for agent clients like Codex, OpenCode, OpenClaw, and CatPaw, and integration notes for coding harnesses including Claude Code, Hermes, and Kilo Code. Maximum output is documented at 128K tokens, which matters for extended agent traces and code generation. Practical fit is best for teams already invested in GUI and tool-calling workflows who want one model to handle reading on-screen interfaces, executing multi-step plans, and producing long structured outputs, while benchmark evidence for GUI and long-process improvements has not yet been published.