Model reference · open weights

NVIDIA-Nemotron-3-Nano

NVIDIA-Nemotron-3-Nano is an open-weight language model from nvidia, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Licence fee required LLMs nvidia 4 variants 877k downloads/mo
Request a licence + hosting quote All served models Not on the shared API today — deployed on request.

About

What NVIDIA-Nemotron-3-Nano is

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Model Overview Model Developer: NVIDIA Corporation Model Dates: September 2025 \- December 2025 Data Freshness: The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Description Nemotron-3-Nano-30B-A3B-BF16 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks. The model employs a hybrid Mixture-of-Experts (MoE) architecture, consisting of 23 Mamba-2 and MoE layers, along with 6 Attention layers. Each MoE layer includes 128 experts plus 1 shared expert, with 6 experts activated per token. The model has 3.5B active parameters and 30B parameters in total. The supported languages include: English, German, Spanish, French, Italian, and Japanese. Improved using Qwen. This model is ready for commercial use. What is Nemotron? NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. To get started, you can use our quickstart guide below. License/Terms of Use Governing Terms: Use of this model is governed by the NVIDIA Nemotron Open Model License. Reasoning Benchmark Evaluations We evaluated our model on the following benchmarks: All evaluation results were collected via Nemo Evaluator SDK and Nemo Skills. The open source container on Nemo Skills packaged via NVIDIA’s Nemo Evaluator SDK used for evaluations can be found here. In addition to Nemo Skills, the evaluations also used dedicated packaged containers for Tau-2 Benc

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makernvidia
TypeLanguage models
Parameters (lead)31.6B
Variants4
Runs withtransformers
Released2025-12-04
Popularity877k downloads / month
Likes817
LicenceCommercial licence needed

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
NVIDIA-Nemotron-3-Nano-30B-A3B-BF1631.6BBF16~72.6 GBWeights ↗
NVIDIA-Nemotron-3-Nano-30B-A3B-FP831.6BFP8~36.3 GBWeights ↗
NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP418.2BNVFP4Weights ↗
NVIDIA-Nemotron-3-Nano-4B-BF164.0BBF16~9.1 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys nvidia-nemotron-3-nano for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (nvidia-nemotron-3-nano below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia-nemotron-3-nano","messages":[{"role":"user","content":"Hello"}]}'

Details

Languages, data & research

Languages

en es fr de ja it

Trained / evaluated on

nvidia/Nemotron-Pretraining-Code-v1 nvidia/Nemotron-CC-v2 nvidia/Nemotron-Pretraining-SFT-v1 nvidia/Nemotron-CC-Math-v1 nvidia/Nemotron-Pretraining-Code-v2 nvidia/Nemotron-Pretraining-Specialized-v1 nvidia/Nemotron-CC-v2.1 nvidia/Nemotron-CC-Code-v1 nvidia/Nemotron-Pretraining-Dataset-sample nvidia/Nemotron-Competitive-Programming-v1 nvidia/Nemotron-Math-v2 nvidia/Nemotron-Agentic-v1 nvidia/Nemotron-Math-Proofs-v1 nvidia/Nemotron-Instruction-Following-Chat-v1

Tags

transformers safetensors nemotron_h text-generation nvidia pytorch conversational custom_code en es fr de ja it

Papers

Licence

Commercial licence needed

The weights are open but its licence needs a commercial agreement for business use. AxForge can arrange that licence and host the model for you — you pay AxForge, we settle with the model’s maker. Ask us for a quote. Read the licence ↗

Sources

Weights & code

Want NVIDIA-Nemotron-3-Nano on EU-owned hardware?

Request a licence + hosting quote See what’s served now

Explore

More language models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms