NVIDIA model product
Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG.
Updated Aug 12, 2026. Default version: Nemotron 3 Ultra (550B A55B)
Technical details for the model's default version.
This profile uses the model's current scored version. Arena ratings and prices are shown separately.
Benchmark scores for Nemotron 3 Ultra (550B A55B).
| Apex | 0.8 | 1 | 2 | 100.0% | C | |
| IMO-AnswerBench | 0.9 | 1 | 19 | 100.0% | C | |
| OmniScience | 0.8 | 1 | 2 | 100.0% | C | |
| PinchBench | 0.9 | 1 | 4 | 100.0% | C | |
| ProfBench | 0.6 | 1 | 1 | 100.0% | C | |
| RULER | 0.9 | 1 | 4 | 100.0% | C | |
| IFBench | 0.8 | 2 | 29 | 96.4% | C | |
| GDPval | 0.5 | 3 | 3 | 0.0% | C | |
| CritPT | 0.0 | 4 | 4 | 0.0% | C | |
| LiveCodeBench v6 | 0.9 | 4 | 53 | 94.2% | C | |
| LongBench v2 | 0.6 | 4 | 17 | 81.3% | C | |
| MMLU-ProX | 0.8 | 5 | 32 | 87.1% | C | |
| TAU3-Bench | 0.2 | 5 | 5 | 0.0% | C | |
| Multi-Challenge | 0.6 | 6 | 29 | 82.1% | C | |
| WMT24++ | 0.8 | 6 | 23 | 77.3% | C | |
| Finance Agent | 0.5 | 8 | 8 | 0.0% | C | |
| AA-LCR | 0.7 | 9 | 16 | 46.7% | C | |
| MMLU-Pro | 0.9 | 9 | 129 | 93.8% | C | |
| SciCode | 0.4 | 9 | 19 | 55.6% | C | |
| Terminal-Bench 2.1 | 0.6 | 17 | 19 | 11.1% | C | |
| Finance Agent v2 | 0.4 | 20 | 26 | 24.0% | B | |
| SWE-bench Multilingual | 0.7 | 21 | 34 | 39.4% | C | |
| Humanity's Last Exam | 0.4 | 37 | 93 | 60.9% | C | |
| GPQA | 0.9 | 42 | 234 | 82.4% | C | |
| BrowseComp | 0.4 | 49 | 58 | 15.8% | C | |
| SWE-Bench Verified | 0.7 | 57 | 105 | 46.1% | C |
Preference and agent-evaluation results for the default version.
| agent tool hallucination | overall | 33 | 0.0 | N/A | 135.6K | |
| agent bash recovery steps | overall | 43 | -0.2 | N/A | 7.8K | |
| agent praise complaint | overall | 43 | -0.1 | N/A | 878 | |
| agent | overall | 44 | -0.1 | N/A | 152.4K | |
| agent steerability | overall | 46 | -0.2 | N/A | 4.6K | |
| agent task outcome explicit | overall | 46 | -0.2 | N/A | 3.6K | |
| text | overall | 49 | 1445.0 | 10,719 | N/A | |
| text factuality | overall | 95 | 1430.0 | 10,719 | N/A | |
| text style control | overall | 96 | 1426.4 | 10,719 | N/A |
Official vendor API pricing appears first, followed by individual provider offers.
| Kenari | nemotron-3-ultra-550b-a55b | global | N/A | N/A | 1M | |
| UnoRouter | nemotron-3-ultra-550b-a55b:free | global | N/A | N/A | 1M | |
| routing.run | nemotron-3-ultra | global | $0.10 | $0.10 | 131.1K | |
| Nvidia | nvidia/nemotron-3-ultra-550b-a55b | global | $0.50 | $2.5 | 1M | |
| NanoGPT | nvidia/nemotron-3-ultra-550b-a55b | global | $0.50 | $2.5 | 1M | |
| LLM Gateway | nemotron-3-ultra-550b | global | $0.50 | $2.5 | 1M | |
| Pioneer | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | global | $0.50 | $2.5 | 1M | |
| Kilo Gateway | nvidia/nemotron-3-ultra-550b-a55b | global | $0.50 | $2.2 | 512.3K | |
| Together AI | nvidia/nemotron-3-ultra-550b-a55b | global | $0.60 | $3.6 | 512.3K | |
| Vercel AI Gateway | nvidia/nemotron-3-ultra-550b-a55b | global | $0.60 | $2.4 | 1M | |
| OpenRouter | nvidia/nemotron-3-ultra-550b-a55b | global | $0.60 | $3.6 | 512.3K |
Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.
Available versions of this model. The score column identifies the version used in the overall ranking.
| Nemotron 3 Ultra (550B A55B) | 62.9 | 550B | 1M | 128K | Yes | OpenMDW License v1.1 |
Key information about Nemotron 3 Ultra and its available data.
Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG.
It uses a hybrid Latent Mixture-of-Experts (LatentMoE) architecture interleaving Mamba-2, MoE, and select Attention layers, with Multi-Token Prediction (MTP) for native speculative decoding, and is pre-trained on ~20T tokens with an NVFP4 recipe.
Reasoning is configurable on/off (plus a medium-effort mode) via the chat template. It supports up to a 1M-token context and 10 languages (English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, Chinese). 1 license.
Data as of 2026-08-11.
Open a comparison with the three ranked models immediately above and below this model.
Recommendations prioritize the same model type and family, then the closest LLMBoard score.
Common questions about Nemotron 3 Ultra.
Nemotron 3 Ultra's default version was released on Jun 4, 2026.
Nemotron 3 Ultra's official API price is $0.50 per million input tokens and $2.5 per million output tokens via Nvidia. The lowest tracked third-party offer starts at $0.10 input and $0.10 output via routing.run.
Nemotron 3 Ultra was created by NVIDIA.
The default version has a 1M token context window.
Yes. The default version is marked as open weight under OpenMDW License v1.1.
11 provider offerings are linked to the default version.
Nearby ranked alternatives include Gemini 3 Flash, MiMo V2 Pro, GPT-5.3-Codex.