Llama 3.3 Nemotron Super 49B (v1): benchmark results
Provider: NVIDIA. Released 2025-03-18. Access: Open.
Unified ELO 1555 ± 31, rank #595 of 1636 rated models, from 83 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RFEval - Context Understanding | 81.7 | Contrast-conditional reasoning faithfulness (%): after a cou | 100 |
| RFEval - Paper Review | 98.47 | Contrast-conditional reasoning faithfulness (%): after a cou | 100 |
| RFEval - Mathematical Reasoning | 44.9 | Contrast-conditional reasoning faithfulness (%): after a cou | 90.9 |
| RFEval | 68.52 | Contrast-conditional reasoning faithfulness (%): after a cou | 90 |
| RFEval - Logical Reasoning | 77.13 | Contrast-conditional reasoning faithfulness (%): after a cou | 81.8 |
| RFEval - Table Reasoning | 69.38 | Contrast-conditional reasoning faithfulness (%): after a cou | 72.7 |
| AA MATH-500 | 95.87 | Accuracy (%) | 70 |
| RFEval - Code Generation | 26.48 | Contrast-conditional reasoning faithfulness (%): after a cou | 68.2 |
| AA Omniscience - Software Engineering (SWE) - HTML | 34 | Accuracy (%) | 65.9 |
| RFEval - Legal Decision | 80.38 | Contrast-conditional reasoning faithfulness (%): after a cou | 63.6 |
| ZeroEval MATH-500 | 96.6 | MATH-500 Score | 62.5 |
| AA MMLU-Pro | 78.46 | Accuracy (%) | 60.6 |
Interactive version: theaggregate.ai/model?slug=llama-3-3-nemotron-super-49b-v1 · How It Works · Data refreshed daily, snapshot 2026-10-11.