AutoBench — leaderboard

Automated multi-metric LLM evaluation ranking models by cost, latency, and domain performance. Models ranked by aggregate score across practical business tasks.

Metric: AutoBench Rank. Source: autobench.org. Status: saturation imminent. 32 models tracked.

Top models

#ModelScore
1GPT-5.14.49
2GPT-54.45
3Gemini 3 Pro (Preview)4.39
4Gemini 2.5 Pro4.37
5GPT-OSS-120B4.37
6Kimi K2 (Thinking)4.34
7Grok 4.1 Fast4.34
8GPT-5 Nano4.32
9Claude Sonnet 4.54.31
10Gemini 2.5 Flash4.3
11Qwen 3 235B A22B (Thinking)4.28
12Claude Haiku 4.54.27
13Claude Opus 4.14.27
14GLM-4.64.25
15Qwen 3 235B A22B 25074.24

Interactive version: theaggregate.ai/benchmark?slug=autobench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.