WildBench — leaderboard

Challenging real-user instruction-following benchmark reporting WildBench scores, task-category scores, Elo, and comparison metrics.

Metric: WB Score Task-Macro. Source: huggingface.co. Status: saturated. 63 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)59.3
2GPT-4o Mini (2024-07-18)57.14
3Mistral Large 2 (Jul)55.57
4Yi Large (Preview)55.29
5GPT-4 Turbo55.22
6Claude 3.5 Sonnet (20240620)54.7
7gemma-2-9B-it-SimPO53.28
8gemma-2-9B-it-DPO53.22
9Gemini 1.5 Pro52.95
10GPT-4 Preview (0125)52.28
11Claude 3 Opus (20240229)51.71
12Yi Large48.93
13Gemini 1.5 Flash48.85
14Gemma 2 27B (IT)48.54
15DeepSeek V2 Chat48.21

Interactive version: theaggregate.ai/benchmark?slug=wildbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.