Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Hy3 (DeepInfra)

Hy3 is Tencent's open-weights release from July 2026, built as a Mixture-of-Experts model with 295 billion total parameters and 21 billion active per pass, paired with native 256K context support. On DeepInfra it is delivered in FP8 quantization, which keeps the large context window practical while reducing memory and compute overhead. The design points clearly at complex reasoning and agentic workloads rather than lightweight chat, with three reasoning modes offered to let callers trade depth against latency on harder prompts.

In Artificial Analysis's cross-provider comparison, DeepInfra's FP8 build is one of four tracked hosts and currently trails the field on raw output throughput, sitting well behind Novita, SiliconFlow FP8, and GMI in tokens per second, and it does not appear among the lowest-latency options. DeepInfra is therefore best chosen when the priorities are FP8 efficiency, broad ecosystem compatibility, and caching-style pricing rather than maximum speed, with the option to switch hosts if throughput becomes the bottleneck. Open weights also make Hy3 a reasonable pick for teams that want to evaluate or self-host the same weights elsewhere while still using DeepInfra as a managed fallback.

LLM Gatewaydeepinfra/hy3Hy

Historical specifications

Powered by
Provider
LLM Gateway
Model key
deepinfra/hy3
Release date
Jul 6, 2026
Last updated
Jul 6, 2026
Input modalities
Output modalities
Capabilities

Historical cost

{
  "input": 0.14,
  "output": 0.58,
  "cache_read": 0.035
}

Historical limits

{
  "input": 192000,
  "output": 131072,
  "context": 262144
}

Latest news about Hy3 (DeepInfra)

LLM Gateway

Coverage

Gulf Business reports on Tencent's international rollout of the open-source Hy3 foundation model, dated 07 August 2026. The article confirms core technical specifications shared with the DeepInfra endpoint: a Mixture-of-Experts architecture with 295 billion total parameters and 21 billion active parameters, support for Beyond the open-source distribution, the report details enterprise deployment channels: API access through Tencent Cloud TokenHub as a model-as-a-service layer with multi-model routing and governance, and integration into Tencent products including WorkBuddy, CodeBuddy, Yuanbao, Marvis, and ima. Tencent states planned

Videos about Hy3 (DeepInfra)

More models around Hy3 (DeepInfra)