AlpacaEval 2.0: leaderboard

Automated instruction-following evaluation using length-controlled win rates against GPT-4 Turbo. 805 prompts scored by an LLM judge. Fast, cheap, and highly correlated with Chatbot Arena.

Metric: LC Win Rate (%). Source: tatsu-lab.github.io. Status: saturated. 223 models tracked.

Top models

#ModelScore
1gemma-2-9B-it-SimPO72.35
2gemma-2-9B-it-DPO67.66
3GPT-4o (2024-05-13)57.46
4GPT-4 Turbo55.02
5Claude 3.5 Sonnet (20240620)52.37
6Yi Large (Preview)51.89
7GPT-4o Mini (2024-07-18)50.73
8Storm-7B50.45
9Llama-3-Instruct-8B-SimPO-ExPO45.78
10Llama-3-Instruct-8B-SimPO44.65
11Qwen 1.5 110B Chat43.91
12Nanbeige2-16B-Chat40.59
13Claude 3 Opus (20240229)40.51
14Infinity-Instruct-7M-Gen-mistral-7B39.67
15Llama 3.1 405B Instruct39.26

Interactive version: theaggregate.ai/benchmark?slug=alpacaeval-2-0 · How It Works · Data refreshed daily, snapshot 2026-09-05.