LLM Releases

Lab release history

Last updated Jul 6, 2026

NVIDIA model releases

GPU platform company increasingly shipping open Nemotron foundation models. This page collects the lab's model releases, lifecycle events, source links, and model metadata in one crawlable record.

7
Models
1
Labs
7
Open
4
Recent

7 models

Nemotron-Labs-3-Puzzle-75B-A9B

Available
NVIDIAOpen weights

A deployment-optimized open-weight model from NVIDIA, released July 6, 2026 — a compressed variant of Nemotron-3-Super-120B-A12B produced with "Iterative Puzzle", a post-training compression framework that jointly prunes MoE experts, active-parameter budget, and Mamba state to boost inference efficiency while preserving accuracy. Reduces the parent from 120.7B total / 12.8B active to 75.3B total / 9.3B active, keeping the hybrid Mamba-Transformer LatentMoE architecture with Multi-Token Prediction. Delivers ~2x higher server throughput than Nemotron-3-Super on a single 8xB200 node at matched user throughput and raises sustainable 1M-token single-H100 concurrency from 1 to 8 requests. Targets collaborative agents, chatbots, RAG, complex instruction-following, and long-context reasoning across English, code, and six other languages. Shipped in BF16, FP8, and NVFP4 variants under the OpenMDW-1.1 license.

Hybrid75.3B1M ctxJul 6, 2026

Nemotron 3 Ultra 550B-A55B

Available
NVIDIAFrontierOpen weights

NVIDIA's largest Nemotron 3 open-weight hybrid Mamba-Transformer MoE, tuned for agentic reasoning, coding, planning, and tool calling.

Hybrid550B1M ctxJun 4, 2026

Nemotron 3 Super 120B-A12B

Available
NVIDIAFrontierOpen weights

Open-weight hybrid Mamba-Transformer MoE designed for collaborative agents and high-volume enterprise workflows.

Hybrid120B1M ctxMar 16, 2026

Nemotron 3 Nano 30B-A3B

Available
NVIDIAOpen weights

Efficient Nemotron 3 MoE checkpoint for agentic reasoning and coding, activating about 3B parameters while supporting 1M-token contexts.

Hybrid30B1M ctxDec 15, 2025

Llama-3.3-Nemotron-Super-49B

Available
NVIDIAOpen weights

Open Llama Nemotron reasoning model from NVIDIA's 2025 Nemotron family.

Dense49B128K ctxApr 2, 2025

Llama-3.1-Nemotron-70B

Available
NVIDIAOpen weights

NVIDIA-tuned Llama 3.1 70B instruction model optimized with Nemotron reward and alignment recipes.

Dense70B128K ctxOct 15, 2024

Nemotron-4 340B

Available
NVIDIAOpen weights

NVIDIA's large open model family for synthetic data generation and reward modeling.

Dense340B4K ctxJun 14, 2024