Hy-MT2-30B-A3B is Tencent's open-source translation model, the second generation of the Hy-MT lineage, designed for mutual translation across 33 languages including Chinese and selected ethnic minority languages. Tencent frames it as Hy Translation 2.0 and reports that it significantly improves instruction-following over the previous generation, handling structured, delimiter-based, contextual, glossary-based, and style-adapted translation tasks that require more than literal word-for-word rendering. The model is positioned for both specialized domains and real business scenarios where translation must obey precise formatting and stylistic rules.
Underneath, the architecture is a sparsely activated Transformer with 30 billion total parameters but only 3 billion active per token, organized across 48 layers where layer 0 is dense and layers 1 through 47 are Mixture-of-Experts. Each MoE layer routes tokens to 8 of 128 routed experts with a sigmoid gating function, plus one shared expert, using Grouped Query Attention with per-head QK RMSNorm and RoPE position encoding. NVIDIA's coverage documents a 256K-token context window, and Tencent cites leading results on open-source translation benchmarks such as Flores200 and WMT25, making the model a strong fit for long-context, multilingual translation workloads where instruction adherence and broad language coverage matter more than general-purpose conversation.