Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities
1,048,576 tokens
Recent tweets and retweets from Deep Infra
Introducing the GOAT plan ๐
Command Code GOAT plan is the best low-cost coding plan on the market today.
- $10/mo to get $70 credits. 7x what you pay
- 30+ open weight and closed models
$1 Go got you started. GOAT gets you shipping.
Subscribe โ be a GOAT.
DeepInfra is serving 700B+ new tokens/day just on @OpenRouter โ and many more for our customers on private endpoints โ all on @nvidia Blackwell Ultra B300s in the US with ZDR ๐บ๐ธ
Open-source models like @deepseek_ai V4 Flash 0731 are production-grade now. More tokens forโฆ
try it here: deepinfra.com/deepseek-ai/Deโฆ
Link
deepseek-ai/DeepSeek-V4-Flash-0731 - Demo - DeepInfra
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.โฆ
๐ Today, weโre releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash.
Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path.
For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput.
The efficiency and accuracy of theโฆ
Ling-3.0-flash is live on DeepInfra
A 124B MoE from @AntLingAGI with just 5.1B active params, large-model capacity at close to small-model cost. Reasoning + tool calling on by default, 128K native context (up to 1M with YaRN).
Built for agents. Live now ๐
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.