Vercel AI Gateway
InclusionAI officially released the weights for Ling-3.0-flash on Hugging Face and ModelScope on August 5, 2026, with both BF16 and FP8 quantized versions available under the MIT license. This followed the initial API announcement in late July and made the 124B-total/5.1B-active MoE model practically downloadable for i Architecture analysis grounded in the official model card confirms the hybrid linear attention design combining Kimi Delta Attention and MLA in a 5:1 ratio, alongside the MoE compute configuration. The piece documents FP8 quantization retention characteristics and discusses adoption considerations for engineering teams