Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) — benchmark results

Llama 3.1 Nemotron Ultra 253B v1 evaluated with reasoning enabled. Provider: NVIDIA. Released 2025-04-07. Access: Open.

Unified ELO 1581 ± 35, rank #529 of 1776 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA MMLU-Pro82.46Accuracy (%)80.8
AA MATH-50095.2Accuracy (%)80.7
AA LiveCodeBench64.13Pass@1 (%)70.5
AA Omniscience - Software Engineering (SWE) - HTML40Accuracy (%)67.5
AA Omniscience - Humanities & Social Sciences22.6Accuracy (%)61.4
AA Omniscience - Software Engineering (SWE) - Dart24Accuracy (%)60.7
AA AIME 202563.67Accuracy (%)59.8
AA GPQA Diamond72.83Accuracy (%)58
AA Omniscience - Software Engineering (SWE) - Julia16Accuracy (%)57.8
AA Omniscience - Law12Accuracy (%)56.1
AA Humanity's Last Exam8.08Accuracy (%)55.3
AA Omniscience - Science, Engineering & Mathematics26.8Accuracy (%)54.2

Interactive version: theaggregate.ai/model?slug=llama-3-1-nemotron-ultra-253b-v1-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.