Skip to content
The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps Task Explorer Builder
How It Works Guesswork Metrics
The Aggregate
Loading data...

ATM-Bench — leaderboard

Metric: Oracle QS (%). Source: atmbench.github.io. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.586
2GPT-585.3
3GPT-5.485.1
4Qwen 3.6 27B84.5
5GPT-5.284.12
6MiniMax-M383.6
7Claude Sonnet 4.583.3
8MiMo-V2.582.6
9Gemini 2.5 Pro78.6
10Qwen 3 VL 8B Instruct78.19
11Gemini 2.5 Flash76.9
12MiMo-V2.5-Pro71.7
13Qwen 3 VL 2B Instruct34.8

Interactive version: theaggregate.ai/benchmark?slug=atm-bench · How It Works · Data refreshed daily, snapshot 2026-08-05.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.