Currently listed through these providers:
Model details
Nemotron 3 Nano 30B A3B FP8
Nemotron 3 Nano 30B A3B FP8 is a quantized FP8 build in NVIDIA's Nemotron family, designed as a single unified model that handles both reasoning and non-reasoning conversational tasks. NVIDIA describes it as a large language model trained from scratch rather than a fine-tune of an existing base, and the accompanying arXiv paper provides the technical lineage behind that design. The FP8 packaging makes the 30B-class active configuration friendlier to deploy on modern accelerator hardware while preserving the behavior of the parent A3B variant, which matters for teams that want reasoning quality without full-precision memory costs.
In practical terms the model fits workloads that need long-context reasoning, tool calling, and configurable sampling for chat-style applications. Its 262K context window, as listed on the Red Hat AI Inference catalog page, supports document-heavy sessions, multi-turn agent loops, and retrieval-augmented pipelines that exceed shorter 32K or 128K budgets. The combination of reasoning and tool-use capability with temperature control makes it well suited to assistant products, code helpers, and analytical agents where a single model can both plan steps and produce final answers.
Quick Info
Powered by- Provider
- Infomaniak
- Model key
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8
- Release date
- Dec 15, 2025
- Last updated
- Aug 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.25
Limits
- Input tokens
- 1,000,000 tokens
- Output tokens
- 262,144 tokens
- Context window
- 1,000,000 tokens