AlpacaEval 1.0 — leaderboard

Original instruction-following evaluation using raw win rates against text-davinci-003. 805 prompts scored by an LLM judge. Predecessor to AlpacaEval 2.0.

Metric: Win Rate (%). Source: tatsu-lab.github.io. Status: saturated. 102 models tracked.

Top models

#ModelScore
1Mistral Medium96.83
2GPT-495.28
3tulu-2-dpo-70B95.03
4Mixtral 8x7B Instruct (v0.1)94.78
5GPT-4 (0314)94.78
6Yi 34B (Chat)94.08
7GPT-4 (0613)93.78
8Mistral 7B Instruct (v0.2)92.78
9Llama 2 70B Chat (HF)92.66
10Claude 291.36
11zephyr-7B-beta90.6
12GPT-3.5 Turbo (0301)89.37
13WizardLM-13B-V1.289.17
14vicuna-33B-v1.388.99
15tulu-2-dpo-13B88.12

Interactive version: theaggregate.ai/benchmark?slug=alpacaeval-1-0 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.