NovitaAI
The KDH News republication of the BusinessWire press release confirms Ant Group's official announcement of Ling-2.6-flash as a sparse Mixture-of-Experts model with 104B total parameters and 7.4B active, designed for token efficiency over benchmark-padding via excessive output. Artificial Analysis figures cited in the r On a hybrid linear MoE stack the model achieves up to 340 tokens per second under 4-card H20 conditions, with Prefill throughput 2.2x that of Nemotron-3-Super, and a stable 215 tokens-per-second output speed placing it in the top tier of its size class. The release states Ling-2.6-flash has been specifically enhanced f
