Ollama Cloud
by Dan Ferguson, Benjamin Crabtree, Abdullahi Olaoye, Timothy Ma, Nirmal Kumar Juluru, Vivian Chen, and Pooja Karadgi on 11 FEB 2026 in Amazon SageMaker AI,...
Model details
Nemotron 3 Nano is designed as one model for reasoning and direct-response workloads. It can produce an internal reasoning trace before its final answer, while a chat-template flag can suppress that trace when a concise response is preferred. This makes it practical for applications that need to balance answer quality, transparency, and response presentation across tasks of varying difficulty.
The model uses a hybrid Mixture-of-Experts design with 23 Mamba-2 and MoE layers plus six attention layers. It has 30 billion total parameters, activates 3.5 billion per token, and selects six experts from 128 plus one shared expert in each MoE layer. An NVFP4 variant is also reported to reach up to four times the throughput on Blackwell B200 through Quantization Aware Distillation, making the architecture especially relevant for efficient inference deployments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Ollama Cloud
by Dan Ferguson, Benjamin Crabtree, Abdullahi Olaoye, Timothy Ma, Nirmal Kumar Juluru, Vivian Chen, and Pooja Karadgi on 11 FEB 2026 in Amazon SageMaker AI,...
Ollama Cloud
Per NVIDIA: “We just launched an ultra-efficient NVFP4 precision version of Nemotron 3 Nano that delivers up to 4x higher throughput on Blackwell B200. Using our new Quantization Aware Distillation method, the NVFP4 ver…