ClinePass
Tahir's Medium architecture breakdown (24 July 2026) describes Kimi K3 as a 2.8-trillion-parameter mixture-of-experts model with 896 total experts, of which only 16 are active per token — an activation ratio of about 1.8%. That ratio is lower than comparable sparse models: Nemotron 3 Ultra activates roughly 4.3%, Mixtr The author emphasizes that with 896 experts spread across an estimated 64–72 GPUs, each token must travel between accelerators to reach its assigned experts, making inter-GPU communication the dominant cost rather than raw FLOPs. The piece situates K3 within a broader narrative of Chinese AI labs reshaping inference ec