AlpacaEval 2.0: leaderboard
Automated instruction-following evaluation using length-controlled win rates against GPT-4 Turbo. 805 prompts scored by an LLM judge. Fast, cheap, and highly correlated with Chatbot Arena.
Metric: LC Win Rate (%). Source: tatsu-lab.github.io. Status: saturated. 223 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gemma-2-9B-it-SimPO | 72.35 |
| 2 | gemma-2-9B-it-DPO | 67.66 |
| 3 | GPT-4o (2024-05-13) | 57.46 |
| 4 | GPT-4 Turbo | 55.02 |
| 5 | Claude 3.5 Sonnet (20240620) | 52.37 |
| 6 | Yi Large (Preview) | 51.89 |
| 7 | GPT-4o Mini (2024-07-18) | 50.73 |
| 8 | Storm-7B | 50.45 |
| 9 | Llama-3-Instruct-8B-SimPO-ExPO | 45.78 |
| 10 | Llama-3-Instruct-8B-SimPO | 44.65 |
| 11 | Qwen 1.5 110B Chat | 43.91 |
| 12 | Nanbeige2-16B-Chat | 40.59 |
| 13 | Claude 3 Opus (20240229) | 40.51 |
| 14 | Infinity-Instruct-7M-Gen-mistral-7B | 39.67 |
| 15 | Llama 3.1 405B Instruct | 39.26 |
Interactive version: theaggregate.ai/benchmark?slug=alpacaeval-2-0 · How It Works · Data refreshed daily, snapshot 2026-09-05.