Claw-Eval-Live - Overall Completion: leaderboard
Metric: Overall completion score (0-100): mean task score over the 105 released Claw-Eval-Live workflow tasks (87 service-backed business workflows, 18 local workspace repairs), each task scored from 0 to 1 by deterministic checks on traces, audit logs, service state and artifacts plus a GPT-5.4 rubric judge for semantic dimensions; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 83.6 |
| 2 | GPT-5.4 | 81.7 |
| 3 | Claude Sonnet 4.6 | 79.9 |
| 4 | GLM-5 | 78.1 |
| 5 | MiniMax-M2.7 | 77.5 |
| 6 | MiMo-V2-Pro | 76.9 |
| 7 | Kimi K2.5 | 76.2 |
| 8 | Gemini 3.1 Pro (Preview) | 74 |
| 9 | Qwen 3.5 397B A17B | 72.7 |
| 10 | Qwen 3.6 Plus | 71.4 |
| 11 | MiniMax-M2.5 | 70.9 |
| 12 | DeepSeek V3.2 | 69.3 |
Interactive version: theaggregate.ai/benchmark?slug=claw-eval-live-overall-completion · How It Works · Data refreshed daily, snapshot 2026-10-07.