tulu-2-dpo-70B — benchmark results
Allen AI's Tulu 2 DPO fine-tune of Llama 2 70B, trained on the Tulu V2 mix and UltraFeedback as an open alternative to Llama 2 Chat. Provider: Allen AI. Released 2023-11-12. Access: Open.
Unified ELO 1416 ± 19, rank #1150 of 1776 rated models, from 29 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AlpacaEval 1.0 | 95.03 | Win Rate (%) | 94.1 |
| JustEval | 4.82 | Avg Score (1-5) | 93.3 |
| RewardBench Chat | 97.49 | Accuracy (%) | 90.1 |
| JustEval - Depth | 4.57 | Score (1-5) | 90 |
| JustEval - Factuality | 4.84 | Score (1-5) | 86.7 |
| JustEval - Helpfulness | 4.85 | Score (1-5) | 80 |
| JustEval - Safety | 4.99 | Score (1-5) | 80 |
| Open CoT - LSAT Analytical Reasoning | 6.09 | CoT Gain (%) | 76.7 |
| BiGGen-Bench | 3.68 | Average Score (1-5) | 76.5 |
| RewardBench | 76.21 | Score (%) | 71 |
| JustEval - Clarity | 4.95 | Score (1-5) | 63.3 |
| JustEval - Engagement | 4.74 | Score (1-5) | 63.3 |
Interactive version: theaggregate.ai/model?slug=tulu-2-dpo-70b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.