Agents' Last Exam — leaderboard

Long-horizon agent benchmark covering economically valuable professional workflows in full operating-system sandboxes, with verifiable task outcomes.

Metric: Pass Rate (%). Source: snorkel.ai. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol30.6
2Claude Opus 4.827
3GPT-5.526.6
4Claude Fable 525.7
5GPT-5.420.5
6Claude Opus 4.720.4
7GLM-5.220.4
8Seed 2.1 Pro19.5
9Gemini 3.1 Pro (Preview)15.8
10DeepSeek V4 Pro12.4
11Qwen 3.7 Max11.8
12GLM-5.111.5
13Kimi K2.69.2
14Qwen 3.6 Plus8.6
15MiMo-V2.58.6

Interactive version: theaggregate.ai/benchmark?slug=agents-last-exam · How the rankings work · Data refreshed daily, snapshot 2026-07-22.