Llama-3.1-Tulu-3-70B-DPO — benchmark results

Ai2's DPO-stage checkpoint of the fully open Tulu 3 post-training recipe, built on Llama 3.1 70B. Provider: Meta. Released 2024-11-20. Access: Open.

Unified ELO 1561 ± 71, rank #589 of 1776 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval82.82Score99.2
Open LLM Leaderboard - MuSR23.4Score98.3
Open LLM Leaderboard - GPQA16.78Score94
Open LLM Leaderboard - MATH Level 544.94Score93.3
Open LLM Leaderboard - BBH45.05Score85.4
Open LLM Leaderboard - MMLU-Pro40.36Score85.3
MATH Level 542.66Accuracy (%)35.2
HREF11.91Average HREF Score (%)33.3
OTIS Mock AIME 2024-254.44Accuracy (%)12.8

Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-70b-dpo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.