Model reference · open weights

NVIDIA-Nemotron-Labs-3-Competitive-Coding

LLMs nvidia Text gen · MoE 1 build Its own licence terms 0 dl/mo

NVIDIA-Nemotron-Labs-3-Competitive-Coding is an open-weight language model from NVIDIA. NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 (NVFP4) weighs 352 GB; it needs more than the reference configurations below.

What it is

Released byNVIDIA
TypeLanguage models
TaskText gen · MoE
Parameters (lead)335.0B
Runs withtransformers
Released2026-09-04
Popularity0 downloads / month
Weights352 GB (NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 (NVFP4), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 (NVFP4)

Weights 352 GB (file size) · runtime overhead from 762 MB on a small card.

How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.

CardThe weights alone
RTX 3060 12 GB … 8× B200 180 GB
16 smaller cards
does not fit

From the model card

What NVIDIA says about NVIDIA-Nemotron-Labs-3-Competitive-Coding

Model Summary

Total Parameters550B (55B active)
ArchitectureBased on Nemotron-3-Ultra
Context LengthUp to 262,144 tokens
Best ForCompetitive programming, algorithmic problem solving, code-reasoning research and benchmarking, test-time-compute research
Reasoning ModeStructured Explanation → Confidence → Answer response format with step-by-step reasoning before the final C++ solution
LicenseOpenMDW License Agreement, version 1.1
Release DateHugging Face: 09/03/2026 via model page

Quick Start

This checkpoint is a competitive-programming specialist model, not a general-purpose chat or agent model. It supports commercial and non-commercial applications, including research and evaluation contexts such as competitive programming benchmarks, code-reasoning research, and test-time-compute studies (for example, GenCorrect-style iterative refinement). See Use Case below.

Model Overview

Model Developer: NVIDIA Corporation

Read the full model card

Model Development: Fine-tuned from NVIDIA-Nemotron-3-Ultra-550B-A55B

What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

Description

Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra, fine-tuned for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 curated problems spanning 16 regional and international competitive-programming contest families. Selected as the SFT teacher for its higher accuracy and roughly 30% shorter generations compared to a DeepSeek-V4-Flash-trained variant, GLM-5.2 distillation yields a model that, combined at inference time with GenCorrect — an iterative closed-loop test-time compute strategy that generates diverse candidate solutions, incorporates evaluator feedback, and refines subsequent generations under a fixed submission budget — was evaluated live and prospectively on the IOI 2026 problem set under official contest time, internet-access, and submission constraints, scoring 535.4 out of 600 and surpassing both the gold-medal threshold (361.12) and the top human contestant's score (498.27), making it the first AI system reported to outscore the highest-scoring human contestant on an IOI problem set.

This model is ready for commercial or non-commercial use.

License/Terms of Use

Governing Download Terms: Use of this model is governed by the OpenMDW License Agreement, version 1.1 (OpenMDW-1.1).

Benchmarks

BenchmarkNemotron-Labs-3-Competitive-Coding
IOI 2025 — with GenCorrect (5 rounds)502.0 / 600
ICPC 2025 — with GenCorrect (5 rounds)9.6 / 12 problems solved
LiveCodeBench Pro — Pass@174.5%
IOI 2026 — live, prospective, competition-specific run535.4 / 600 (Gold; exceeds gold threshold of 361.12 and top human score of 498.27)

All results are from the source paper, Post-Training Language Models for Gold-Medal Performance in Coding Competitions (NVIDIA, arXiv:2609.02849). IOI Score@1/Score@200 and GenCorrect results are averaged over multiple independent runs; see the paper for full methodology. IOI 2025 was used as a development benchmark; IOI 2026 results are from a single prospective live run conducted under official IOI time, internet-access, and submission constraints before problems were publicly released, and were not part of the official IOI rankings.

Deployment Geography: Global

Use Case

Nemotron-Labs-3-Competitive-Coding is intended for researchers and developers evaluating or advancing frontier code-reasoning capability, particularly on competitive programming and algorithmic problem solving where a solution must satisfy strict correctness, efficiency, and hidden test-case constraints. It is suited to benchmarking and research on long-horizon reasoning, agentic code generation, and test-time compute strategies such as GenCorrect-style iterative refinement, rather than general-purpose chat, instruction-following, or production coding-assistant deployment. Use in safety-critical or real-time production systems requires further evaluation and safeguards appropriate to the application.

Release Date

Hugging Face: 09/03/2026 via model page

Reference(s)

Model Architecture

  • Architecture Type: Other
  • Architecture Description: Mamba2-Transformer Hybrid Latent Mixture of Experts (LatentMoE) with Multi-Token Prediction (MTP)
  • Network Architecture: Other
  • Network Architecture Description: Nemotron Hybrid LatentMoE
  • Base Model: Nemotron-3-Ultra-550B-A55B
  • Number of model parameters: 550B Total / 55B Active

This model inherits its architecture unchanged from Nemotron-3-Ultra. See the Nemotron-3-Ultra Technical Report for architecture details.

Input

Input Type(s): Text

Input Format(s): String

Input Parameters: One-Dimensional (1D)

Other Properties Related to Input: Maximum context length up to 262,144 tokens

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms