ALL Bench LLM — leaderboard
Composite LLM leaderboard aggregating cross-verified scores across reasoning, knowledge, coding, and instruction-following evaluations.
Metric: Average Numeric Benchmark Score (%). Source: huggingface.co. Status: saturation imminent. 39 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.1 Fast | 81.77 |
| 2 | Step 3.5 Flash | 77.28 |
| 3 | Claude Sonnet 4.6 | 75.9 |
| 4 | qwen3.5-flash | 75 |
| 5 | GPT-5.3 Codex | 74.67 |
| 6 | Gemini 3.1 Pro (Preview) | 74.62 |
| 7 | Qwen 3.5 397B A17B | 73.94 |
| 8 | Claude Opus 4.6 | 73.69 |
| 9 | GPT-5.2 | 73.02 |
| 10 | Gemini 3 Flash | 72.09 |
| 11 | Claude Sonnet 4.5 | 72.07 |
| 12 | Qwen 3.5 4B | 71.22 |
| 13 | Phi-4 | 71 |
| 14 | MiniMax-M2.5 | 69.47 |
| 15 | Gemini 3 Pro | 69.45 |
Interactive version: theaggregate.ai/benchmark?slug=all-bench-llm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.