Llama-3.1-Tulu-3-70B-DPO — benchmark results
Ai2's DPO-stage checkpoint of the fully open Tulu 3 post-training recipe, built on Llama 3.1 70B. Provider: Meta. Released 2024-11-20. Access: Open.
Unified ELO 1561 ± 71, rank #589 of 1776 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 82.82 | Score | 99.2 |
| Open LLM Leaderboard - MuSR | 23.4 | Score | 98.3 |
| Open LLM Leaderboard - GPQA | 16.78 | Score | 94 |
| Open LLM Leaderboard - MATH Level 5 | 44.94 | Score | 93.3 |
| Open LLM Leaderboard - BBH | 45.05 | Score | 85.4 |
| Open LLM Leaderboard - MMLU-Pro | 40.36 | Score | 85.3 |
| MATH Level 5 | 42.66 | Accuracy (%) | 35.2 |
| HREF | 11.91 | Average HREF Score (%) | 33.3 |
| OTIS Mock AIME 2024-25 | 4.44 | Accuracy (%) | 12.8 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-70b-dpo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.