Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Last updated
Apr 20, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities
262,144 tokens
Recent tweets and retweets from Deep Infra
Details 👇
deepinfra.com/blog/flex-serv…
Link
Introducing the Flex Service Tier: Cheaper Inference When You Can Wait
Low pay-as-you-go pricing. No long-term contracts. Simple APIs. Scale to trillions of tokens. 100+ AI models.
deepinfra.com
🆕New on DeepInfra: the Flex service tier.
Not every request needs an answer this second. Flex runs latency-tolerant work; evals, data enrichment, offline jobs at 0.8× the real-time price (20% off).
One field: service_tier="flex". And if it's not served, it's not billed.
1B BF16
deepinfra.com/nvidia/Nemotro…
Link
nvidia/Nemotron-3-Embed-1B-BF16 - Demo - DeepInfra
Nemotron-3-Embed-1B-BF16 is a compact multilingual text embedding model from NVIDIA, pruned and distilled from Ministral-3 to ~1B parameters, that maps text into 2048-dimensional…
1B NVFP4
deepinfra.com/nvidia/Nemotro…
Link
nvidia/Nemotron-3-Embed-1B-NVFP4 - Demo - DeepInfra
Nemotron-3-Embed-1B-NVFP4 is the NVFP4-quantized version of Nemotron-3-Embed-1B-BF16 — a multilingual text embedding model from NVIDIA that maps text into 2048-dimensional dense…
8B Frontier Accuracy deepinfra.com/nvidia/Nemotro…
Link
nvidia/Nemotron-3-Embed-8B - Demo - DeepInfra
Nemotron-3-Embed-8B is a multilingual text embedding model from NVIDIA, based on Ministral-3-8B, that maps text into 4096-dimensional dense vectors for retrieval and…
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.