AIHubMix
Meituan released LongCat-2.0, a large-scale Mixture-of-Experts language model with 1.6 trillion total parameters and roughly 48 billion activated per token, designed around agentic coding tasks such as code understanding, generation, and execution inside agent workflows. The model supports a native 1-million-token cont The architecture is built around four cost-reducing ideas, with the headline feature being LongCat Sparse Attention — which combines Streaming-aware Indexing, Cross-Layer Indexing, and Hierarchical Indexing to convert fragmented memory access into coalesced HBM reads, amortize indexing across adjacent layers, and apply