tulu-2-dpo-7B — benchmark results
Provider: Allen AI. Released 2023-11-12. Access: Open.
Unified ELO 1319 ± 21, rank #1495 of 1776 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Chat | 97.49 | Accuracy (%) | 90.1 |
| RewardBench | 72.12 | Score (%) | 60.2 |
| JustEval - Depth | 4.36 | Score (1-5) | 60 |
| AlpacaEval 1.0 | 84.22 | Win Rate (%) | 54.5 |
| JustEval - Engagement | 4.69 | Score (1-5) | 46.7 |
| JustEval - Helpfulness | 4.64 | Score (1-5) | 46.7 |
| BiGGen-Bench | 3.28 | Average Score (1-5) | 44.1 |
| JustEval | 4.67 | Avg Score (1-5) | 36.7 |
| JustEval - Clarity | 4.92 | Score (1-5) | 36.7 |
| RewardBench Chat Hard | 56.14 | Accuracy (%) | 36.1 |
| Open CoT - LSAT Analytical Reasoning | 2.61 | CoT Gain (%) | 34 |
| JustEval - Factuality | 4.53 | Score (1-5) | 33.3 |
Interactive version: theaggregate.ai/model?slug=tulu-2-dpo-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.