Llama 3.1 Nemotron 70B Instruct — benchmark results

NVIDIA's RLHF-tuned Llama 3.1 70B (HelpSteer2-Preference + REINFORCE), topping Arena Hard and AlpacaEval 2 at release (October 2024). Provider: NVIDIA. Released 2024-10-15. Access: Open.

Unified ELO 1533 ± 13, rank #671 of 1776 rated models, from 194 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Portuguese68.25Average Score (%)97.7
Open Japanese LLM - MMLU EN Exact Match81.23Score (%)97.6
Open Japanese LLM - Jsick Exact Match86.62Score (%)97.1
Open CoT - LSAT Logical Reasoning21.96CoT Gain (%)96.9
Open PL LLM - Generative69.13Average Generative Score (%)96.9
StickToYourRole80.66Cardinal Score96.8
EuroEval Finnish NLU - Scandisent FI92.66Sentiment classification Score (%)94.1
Open Japanese LLM - Wiki Coreference SET F19.84Score (%)94
EuroEval Spanish NLU - Sentiment Headlines ES51.12Sentiment classification Score (%)93.9
EuroEval Danish Knowledge91.41Knowledge Average Score (%)92.8
Open CoT - LSAT Analytical Reasoning8.7CoT Gain (%)92.4
Open LLM Leaderboard - MATH Level 542.67Score92

Interactive version: theaggregate.ai/model?slug=llama-3-1-nemotron-70b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.