Agents' Last Exam (OpenClaw): leaderboard
Metric: Overall Pass Rate (%). Source: snorkel.ai. Saturation forecast: Rough model projection: around 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (ALE Run: harness=OpenClaw; effort=High; source-model=GPT-5.5) | 21.71 |
| 2 | GPT-5.4 (ALE Run: harness=OpenClaw; effort=High; source-model=GPT-5.4) | 20.5 |
| 3 | Claude Opus 4.7 (ALE Run: harness=OpenClaw; effort=High; source-model=Claude%20Opus%204.7) | 15.79 |
| 4 | Gemini 3.1 Pro (ALE Run: harness=OpenClaw; effort=High; source-model=Gemini%203.1%20Pro) | 14.14 |
| 5 | Seed 2.1 Pro (ALE Run: harness=OpenClaw; effort=High; source-model=Seed%202.1%20Pro) | 13.1 |
| 6 | DeepSeek V4 Pro (ALE Run: harness=OpenClaw; effort=High; source-model=DeepSeek%20V4%20Pro) | 12.39 |
| 7 | Qwen3.7-Max (ALE Run: harness=OpenClaw; effort=High; source-model=Qwen3%207-Max) | 11.84 |
| 8 | GLM-5.1 (ALE Run: harness=OpenClaw; effort=High; source-model=GLM-5.1) | 11.51 |
| 9 | Kimi K2.6 (ALE Run: harness=OpenClaw; effort=High; source-model=Kimi%20K2.6) | 9.21 |
| 10 | Qwen3.6-Plus (ALE Run: harness=OpenClaw; effort=High; source-model=Qwen3%206-Plus) | 8.55 |
Interactive version: theaggregate.ai/benchmark?slug=agents-last-exam-openclaw · How It Works · Data refreshed daily, snapshot 2026-10-09.