tulu-2-dpo-70B — benchmark results

Allen AI's Tulu 2 DPO fine-tune of Llama 2 70B, trained on the Tulu V2 mix and UltraFeedback as an open alternative to Llama 2 Chat. Provider: Allen AI. Released 2023-11-12. Access: Open.

Unified ELO 1416 ± 19, rank #1150 of 1776 rated models, from 29 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AlpacaEval 1.095.03Win Rate (%)94.1
JustEval4.82Avg Score (1-5)93.3
RewardBench Chat97.49Accuracy (%)90.1
JustEval - Depth4.57Score (1-5)90
JustEval - Factuality4.84Score (1-5)86.7
JustEval - Helpfulness4.85Score (1-5)80
JustEval - Safety4.99Score (1-5)80
Open CoT - LSAT Analytical Reasoning6.09CoT Gain (%)76.7
BiGGen-Bench3.68Average Score (1-5)76.5
RewardBench76.21Score (%)71
JustEval - Clarity4.95Score (1-5)63.3
JustEval - Engagement4.74Score (1-5)63.3

Interactive version: theaggregate.ai/model?slug=tulu-2-dpo-70b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.