Currently listed through these providers:
Model details
Nemotron 3 Nano Omni 30B A3B Reasoning
Built on a hybrid reasoning Mixture-of-Experts design with around 30 billion total parameters while activating roughly 3 billion per token, this model focuses on agentic workflows that mix language, audio, video, images, and documents into a single text response. The architecture targets efficiency: the sparse activation keeps per-request compute modest while the multimodal encoder stack allows users to drop in cross-format context without separate pipelines. Unsloth's documentation highlights its day-zero support from NVIDIA and emphasizes that it is the most capable omnidirectional model in its size class and the most efficient open multimodal release, although those positioning claims come from partner marketing copy rather than independent benchmark validation.
In practice, the open weights invite local experimentation through community GGUF quantizations that fit on roughly 25 GB of memory at four bits and around 36 GB at eight bits, making the system runnable on a single high-memory machine. NVIDIA separates two serving modes, a higher-temperature Thinking configuration for chain-of-thought reasoning and a tighter Instruct configuration for direct answers, so users can tune verbosity without retraining. The combination of a long context window, broad modality coverage, and open release makes the model a sensible choice for teams building grounded assistants, document-vision tools, or audio-aware agents that need flexible deployment beyond a closed API.
Quick Info
Powered by- Provider
- Crusoe
- Model key
- nvidia/Nemotron-3-Nano-Omni-Reasoning-30B-A3B
- Release date
- Apr 28, 2026
- Last updated
- Apr 28, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $1.83
Limits
- Output tokens
- 65,536 tokens
- Context window
- 256,000 tokens
Transparent token rates
Compare Nemotron 3 Nano Omni 30B A3B Reasoning pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Nemotron 3 Nano Omni 30B A3B Reasoning
No articles yet. Fetch the latest news to show it here.