Llama 3 70B Orpo V0.1: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1559 ± 21, rank #839 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - GSM8K76.8Accuracy (%) (5-shot)99.4
Open LLM Leaderboard v1 - MMLU79.39Accuracy (%) (5-shot)99.4
Open LLM Leaderboard v1 - WinoGrande85.48Accuracy (%) (5-shot)98.4
Open LLM Leaderboard v1 - HellaSwag88.01Normalized accuracy (%) (10-shot)89.2
Open LLM Leaderboard - MuSR45.34Score85
Open LLM Leaderboard v1 - ARC Challenge68.69Normalized accuracy (%) (25-shot)81.1
Open LLM Leaderboard - MMLU-Pro38.93Score70.2
Open LLM Leaderboard - MATH Level 515.79Score63.4
Open LLM Leaderboard v1 - TruthfulQA MC249.62MC2 (%) (0-shot)45.5
Open LLM Leaderboard - BBH46.55Score37.6
Open LLM Leaderboard - IFEval20.49Score12.7
Open LLM Leaderboard - GPQA25.76Score9.7

Interactive version: theaggregate.ai/model?slug=llama-3-70b-orpo-v0-1 · How It Works · Data refreshed daily, snapshot 2026-09-23.