Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Last updated
Jul 30, 2026
Input modalities
Output modalities
Capabilities
524,288 tokens
Recent tweets and retweets from Deep Infra
Introducing the GOAT plan ๐
Command Code GOAT plan is the best low-cost coding plan on the market today.
- $10/mo to get $70 credits. 7x what you pay
- 30+ open weight and closed models
$1 Go got you started. GOAT gets you shipping.
Subscribe โ be a GOAT.
DeepInfra is serving 700B+ new tokens/day just on @OpenRouter โ and many more for our customers on private endpoints โ all on @nvidia Blackwell Ultra B300s in the US with ZDR ๐บ๐ธ
Open-source models like @deepseek_ai V4 Flash 0731 are production-grade now. More tokens forโฆ
try it here: deepinfra.com/deepseek-ai/Deโฆ
Link
deepseek-ai/DeepSeek-V4-Flash-0731 - Demo - DeepInfra
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.โฆ
๐ Today, weโre releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash.
Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path.
For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput.
The efficiency and accuracy of theโฆ
Ling-3.0-flash is live on DeepInfra
A 124B MoE from @AntLingAGI with just 5.1B active params, large-model capacity at close to small-model cost. Reasoning + tool calling on by default, 128K native context (up to 1M with YaRN).
Built for agents. Live now ๐
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.