Model details
Meituan/Longcat-Flash-Chat
LongCat-Flash-Chat is a large-scale Mixture-of-Experts model featuring 560 billion total parameters, designed to balance broad knowledge with computational efficiency. By utilizing a zero-computation expert gating mechanism, the model dynamically activates an average of 27 billion parameters per token, allowing it to maintain high throughput without sacrificing performance. This architecture is specifically optimized for conversational and agentic tasks, prioritizing immediate, direct responses over internal reasoning chains. Its design focuses on stability and consistency, making it a reliable choice for applications that require structured function calls and the maintenance of task state across long, multi-turn interactions.
The model benefits from advanced scaling strategies, including hyperparameter transfer, deterministic computation, and multi-stage optimization, which ensure training stability across its massive parameter breadth. These methods support its practical strengths in instruction following, coding, and complex tool-augmented workflows. By delivering competitive performance on benchmarks like MMLU and ArenaHard-V2 while activating fewer parameters than its total size suggests, the model serves as a robust foundation for developers building real-time applications. Its ability to handle long-context inputs while keeping response formatting stable positions it as a strong candidate for automated agents that must execute sequential steps without behavior drift.
Quick Info
Powered by- Provider
- Qiniu
- Model key
- meituan/longcat-flash-chat
- Release date
- Nov 5, 2025
- Last updated
- Nov 5, 2025
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens
Latest news about Meituan/Longcat-Flash-Chat
No articles yet. Fetch the latest news to show it here.