Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities
1,048,576 tokens
Recent tweets and retweets from Deep Infra
🆕The gap between prototyping an AI agent and running one reliably, around the clock, has always been infrastructure: provisioning a VM, securing SSH, wiring up model access, maintaining updates and backups.
We built Hosted Agents to remove that gap. Choose a framework and a…
This is the one we've been excited about. Vera is the CPU built for exactly what we do — running agents at scale. Thanks for featuring the results, @NVIDIAAIInfra
Proud to be the independent benchmark behind @nvidia Vera launch.
We run agents in production so we measured the CPU built for them. 2× faster than the best x86 on agentic orchestration, more agents per core.
Nikola Borisov bet the whole company on a boring word: inference. Deep Infra now runs 5 trillion tokens a week.
Our cover story on @DeepInfra
yespress.io/nikola-borisov?u… via Yespress
Link
Nikola Borisov
The co-founder who wagered Deep Infra on the unsexy half of AI:…
🆕Tencent Hy3 is now on DeepInfra.
A 295B MoE model (21B active) that punches way above its weight. Tencent says it rivals open models with 2–5× the parameters.
$0.14 in / $0.58 out per 1M tokens. 256K context. Function calling + JSON.
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.