Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Infomaniak logo

Model details

Nemotron 3 Nano 30B A3B FP8

Nemotron 3 Nano 30B A3B FP8 is a quantized FP8 build in NVIDIA's Nemotron family, designed as a single unified model that handles both reasoning and non-reasoning conversational tasks. NVIDIA describes it as a large language model trained from scratch rather than a fine-tune of an existing base, and the accompanying arXiv paper provides the technical lineage behind that design. The FP8 packaging makes the 30B-class active configuration friendlier to deploy on modern accelerator hardware while preserving the behavior of the parent A3B variant, which matters for teams that want reasoning quality without full-precision memory costs.

In practical terms the model fits workloads that need long-context reasoning, tool calling, and configurable sampling for chat-style applications. Its 262K context window, as listed on the Red Hat AI Inference catalog page, supports document-heavy sessions, multi-turn agent loops, and retrieval-augmented pipelines that exceed shorter 32K or 128K budgets. The combination of reasoning and tool-use capability with temperature control makes it well suited to assistant products, code helpers, and analytical agents where a single model can both plan steps and produce final answers.

Infomaniaknvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8nemotronbeta

Quick Info

Powered by
Provider
Infomaniak
Model key
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8
Release date
Dec 15, 2025
Last updated
Aug 1, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.25

Limits

Input tokens
1,000,000 tokens
Output tokens
262,144 tokens
Context window
1,000,000 tokens

Latest news about Nemotron 3 Nano 30B A3B FP8

Videos about Nemotron 3 Nano 30B A3B FP8

More models around Nemotron 3 Nano 30B A3B FP8