LiveBench — leaderboard

Continuously updated benchmark measuring many capabilities. Sort models by weighted score or sub-task scores. Tests math, coding, reasoning, language, instruction following, and data analysis.

Metric: LiveBench average (self-reported). Source: benchmarklist.com. Status: saturation imminent. 43 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)81.28
2GPT-5.4 (xHigh)80.91
3Gemini 3.1 Pro (Preview) (High)80.71
4Claude Opus 4.8 (xHigh)77.55
5Claude Opus 4.7 (xHigh)77.1
6Claude Opus 4.6 (Thinking, High)76.79
7Claude Opus 4.8 (High)76.18
8Gemini 3.5 Flash (High)75.79
9Claude Fable 5 (High)75.74
10Claude Sonnet 4.6 (Thinking, High)75.59
11GPT-5.2 (High)75.38
12Qwen 3.7 Max (Max)75.15
13Claude Opus 4.7 (High)74.66
14DeepSeek V4 Pro74.39
15GPT-5.2 Codex74.33

Interactive version: theaggregate.ai/benchmark?slug=livebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.