Hy3 preview is a 295-billion-parameter mixture-of-experts language model from Tencent, released as an open-weight checkpoint on Hugging Face under the tencent/Hy3-preview identifier. Its architecture stacks 80 transformer layers, beginning with a single dense layer followed by 79 MoE blocks, each containing 192 routed experts alongside one shared expert selected through top-8 sigmoid routing. Attention is handled by Grouped Query Attention with 64 query heads and 8 key-value heads at a head dimension of 128, augmented with per-head QK RMSNorm applied before rotary position embeddings, while an expert bias buffer feeds an e_score_correction_bias gate that helps balance expert load during inference. A long-context RoPE configuration with a theta of roughly 11.16 million extends the usable window to around 256,000 tokens, a figure that surfaces as 262K on public routes.
The design intent behind Hy3 preview is production-oriented agentic use, pairing the long context with configurable reasoning effort that can be switched between disabled, low, and high modes so callers can trade latency for depth on a per-request basis. Vendor descriptions emphasize reliable multi-step workflow execution and strong code generation, framing the model as a high-efficiency MoE suitable for tool-using agents and real-world pipelines rather than as a research artifact. For practitioners, the combination of open weights, an exceptionally wide context window, code-focused training emphasis, and adjustable reasoning levels makes it a natural fit for long-document assistants, multi-file code tasks, and agent loops that need sustained state across hundreds of thousands of tokens.