Agents' Last Exam: leaderboard

Long-horizon agent benchmark covering economically valuable professional workflows in full operating-system sandboxes, with verifiable task outcomes.

Metric: Pass Rate (%). Source: snorkel.ai. Status: saturation imminent. 25 models tracked.

Top models

#ModelScore
1Claude Opus 531.6
2GPT-5.6 Sol30.6
3GPT-5.6 Luna30.3
4Kimi K328.3
5GPT-5.6 Terra28
6Claude Opus 4.827
7Grok 4.527
8GPT-5.526.6
9Claude Fable 525.7
10Claude Opus 4.721.1
11GPT-5.420.5
12GLM-5.220.4
13Composer 2.520.4
14Seed 2.1 Pro19.5
15Gemini 3.1 Pro (Preview)16.4

Interactive version: theaggregate.ai/benchmark?slug=agents-last-exam · How It Works · Data refreshed daily, snapshot 2026-09-05.