Tencent's Hy3 is a Mixture-of-Experts large language model built by the Hunyuan team, with 295 billion total parameters arranged across 192 experts but only 21 billion active per token through top-8 routing. This sparse design lets the model deliver near-frontier capability while keeping compute per request far lower than denser peers. A 3.8-billion-parameter multi-token prediction layer enables speculative decoding, pushing throughput to roughly 201 tokens per second. The weights are released under Apache 2.0 with no regional or field-of-use restrictions, supporting full commercial reuse. Hy3 ships with a 256K token context window and was finalized on July 6, 2026, upgrading a preview version that had been available since April 23, 2026.
A defining feature of Hy3 is its hybrid fast-and-slow thinking approach, in which the model decides per request how much reasoning effort to spend, balancing simple queries against tasks that benefit from deeper deliberation. The architecture leans toward efficiency and cost rather than maximum coding or context length, positioning the model as a practical middle ground for production workloads such as coding agents, retrieval-augmented assistants, and structured-output pipelines. Independent write-ups describe it as a strong value play for sub-150K-token sessions, where the combination of sparse activation, open licensing, and adaptive reasoning offers a sensible tradeoff between capability, reliability, and operating cost.