Claw-Eval-Live - Overall Completion: leaderboard

Metric: Overall completion score (0-100): mean task score over the 105 released Claw-Eval-Live workflow tasks (87 service-backed business workflows, 18 local workspace repairs), each task scored from 0 to 1 by deterministic checks on traces, audit logs, service state and artifacts plus a GPT-5.4 rubric judge for semantic dimensions; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.683.6
2GPT-5.481.7
3Claude Sonnet 4.679.9
4GLM-578.1
5MiniMax-M2.777.5
6MiMo-V2-Pro76.9
7Kimi K2.576.2
8Gemini 3.1 Pro (Preview)74
9Qwen 3.5 397B A17B72.7
10Qwen 3.6 Plus71.4
11MiniMax-M2.570.9
12DeepSeek V3.269.3

Interactive version: theaggregate.ai/benchmark?slug=claw-eval-live-overall-completion · How It Works · Data refreshed daily, snapshot 2026-10-07.