AlpacaEval 2.0 — leaderboard
Automated instruction-following evaluation using length-controlled win rates against GPT-4 Turbo. 805 prompts scored by an LLM judge. Fast, cheap, and highly correlated with Chatbot Arena.
Metric: LC Win Rate (%). Source: tatsu-lab.github.io. Status: saturation imminent. 223 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gemma-2-9B-it-WPO-HB | 76.73 |
| 2 | gemma-2-9B-it-SimPO | 72.35 |
| 3 | gemma-2-9B-it-DPO | 67.66 |
| 4 | FuseChat-Llama-3.1-8B-Instruct | 65.39 |
| 5 | FuseChat-Qwen-2.5-7B-Instruct | 63.58 |
| 6 | GPT-4o (2024-05-13) | 57.46 |
| 7 | GPT-4 Turbo | 55.02 |
| 8 | FuseChat-Llama-3.2-3B-Instruct | 54 |
| 9 | Claude 3.5 Sonnet (20240620) | 52.37 |
| 10 | Yi Large (Preview) | 51.89 |
| 11 | GPT-4o Mini (2024-07-18) | 50.73 |
| 12 | Storm-7B | 50.45 |
| 13 | Infinity-Instruct-7M-Gen-Llama3.1-70B | 46.1 |
| 14 | Llama-3-Instruct-8B-SimPO-ExPO | 45.78 |
| 15 | Llama-3-Instruct-8B-SimPO | 44.65 |
Interactive version: theaggregate.ai/benchmark?slug=alpacaeval-2-0 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.