Agents' Last Exam (OpenClaw): leaderboard

Metric: Overall Pass Rate (%). Source: snorkel.ai. Saturation forecast: Rough model projection: around 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.5 (ALE Run: harness=OpenClaw; effort=High; source-model=GPT-5.5)21.71
2GPT-5.4 (ALE Run: harness=OpenClaw; effort=High; source-model=GPT-5.4)20.5
3Claude Opus 4.7 (ALE Run: harness=OpenClaw; effort=High; source-model=Claude%20Opus%204.7)15.79
4Gemini 3.1 Pro (ALE Run: harness=OpenClaw; effort=High; source-model=Gemini%203.1%20Pro)14.14
5Seed 2.1 Pro (ALE Run: harness=OpenClaw; effort=High; source-model=Seed%202.1%20Pro)13.1
6DeepSeek V4 Pro (ALE Run: harness=OpenClaw; effort=High; source-model=DeepSeek%20V4%20Pro)12.39
7Qwen3.7-Max (ALE Run: harness=OpenClaw; effort=High; source-model=Qwen3%207-Max)11.84
8GLM-5.1 (ALE Run: harness=OpenClaw; effort=High; source-model=GLM-5.1)11.51
9Kimi K2.6 (ALE Run: harness=OpenClaw; effort=High; source-model=Kimi%20K2.6)9.21
10Qwen3.6-Plus (ALE Run: harness=OpenClaw; effort=High; source-model=Qwen3%206-Plus)8.55

Interactive version: theaggregate.ai/benchmark?slug=agents-last-exam-openclaw · How It Works · Data refreshed daily, snapshot 2026-10-09.