ALL Bench LLM — leaderboard

Composite LLM leaderboard aggregating cross-verified scores across reasoning, knowledge, coding, and instruction-following evaluations.

Metric: Average Numeric Benchmark Score (%). Source: huggingface.co. Status: saturation imminent. 39 models tracked.

Top models

#ModelScore
1Grok 4.1 Fast81.77
2Step 3.5 Flash77.28
3Claude Sonnet 4.675.9
4qwen3.5-flash75
5GPT-5.3 Codex74.67
6Gemini 3.1 Pro (Preview)74.62
7Qwen 3.5 397B A17B73.94
8Claude Opus 4.673.69
9GPT-5.273.02
10Gemini 3 Flash72.09
11Claude Sonnet 4.572.07
12Qwen 3.5 4B71.22
13Phi-471
14MiniMax-M2.569.47
15Gemini 3 Pro69.45

Interactive version: theaggregate.ai/benchmark?slug=all-bench-llm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.