AlpacaEval 2.0 — leaderboard

Automated instruction-following evaluation using length-controlled win rates against GPT-4 Turbo. 805 prompts scored by an LLM judge. Fast, cheap, and highly correlated with Chatbot Arena.

Metric: LC Win Rate (%). Source: tatsu-lab.github.io. Status: saturation imminent. 223 models tracked.

Top models

#ModelScore
1gemma-2-9B-it-WPO-HB76.73
2gemma-2-9B-it-SimPO72.35
3gemma-2-9B-it-DPO67.66
4FuseChat-Llama-3.1-8B-Instruct65.39
5FuseChat-Qwen-2.5-7B-Instruct63.58
6GPT-4o (2024-05-13)57.46
7GPT-4 Turbo55.02
8FuseChat-Llama-3.2-3B-Instruct54
9Claude 3.5 Sonnet (20240620)52.37
10Yi Large (Preview)51.89
11GPT-4o Mini (2024-07-18)50.73
12Storm-7B50.45
13Infinity-Instruct-7M-Gen-Llama3.1-70B46.1
14Llama-3-Instruct-8B-SimPO-ExPO45.78
15Llama-3-Instruct-8B-SimPO44.65

Interactive version: theaggregate.ai/benchmark?slug=alpacaeval-2-0 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.