Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Last updated
Apr 1, 2026
Input modalities
Output modalities
Capabilities
262,144 tokens
Recent tweets and retweets from Deep Infra
Details 👇
deepinfra.com/blog/flex-serv…
Link
Introducing the Flex Service Tier: Cheaper Inference When You Can Wait
Low pay-as-you-go pricing. No long-term contracts. Simple APIs. Scale to trillions of tokens. 100+ AI models.
deepinfra.com
🆕New on DeepInfra: the Flex service tier.
Not every request needs an answer this second. Flex runs latency-tolerant work; evals, data enrichment, offline jobs at 0.8× the real-time price (20% off).
One field: service_tier="flex". And if it's not served, it's not billed.
1B BF16
deepinfra.com/nvidia/Nemotro…
Link
nvidia/Nemotron-3-Embed-1B-BF16 - Demo - DeepInfra
Nemotron-3-Embed-1B-BF16 is a compact multilingual text embedding model from NVIDIA, pruned and distilled from Ministral-3 to ~1B parameters, that maps text into 2048-dimensional…
1B NVFP4
deepinfra.com/nvidia/Nemotro…
Link
nvidia/Nemotron-3-Embed-1B-NVFP4 - Demo - DeepInfra
Nemotron-3-Embed-1B-NVFP4 is the NVFP4-quantized version of Nemotron-3-Embed-1B-BF16 — a multilingual text embedding model from NVIDIA that maps text into 2048-dimensional dense…
8B Frontier Accuracy deepinfra.com/nvidia/Nemotro…
Link
nvidia/Nemotron-3-Embed-8B - Demo - DeepInfra
Nemotron-3-Embed-8B is a multilingual text embedding model from NVIDIA, based on Ministral-3-8B, that maps text into 4096-dimensional dense vectors for retrieval and…
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.