BenchLM — leaderboard

Composite LLM leaderboard aggregating current model performance across agentic, coding, reasoning, grounded multimodal, knowledge, multilingual, instruction-following, and math categories.

Metric: Overall Score. Source: benchlm.ai. Status: saturation imminent. 200 models tracked.

Top models

#ModelScore
1Claude Mythos 583.9
2Claude Fable 583.7
3GPT-5.6 Sol82
4Kimi K381
5Claude Opus 4.878.3
6Muse Spark 1.177.4
7Grok 4.576.7
8GPT-5.474.2
9GPT-5.573.5
10Qwen 3.7 Max72.8
11GPT-5.6 Terra72.6
12Claude Opus 4.771.9
13Muse Spark71
14MiMo-V2.5-Pro70.2
15MiniMax-M369.8

Interactive version: theaggregate.ai/benchmark?slug=benchlm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.