Claw-Eval-Live — leaderboard

Quarterly refreshed enterprise-workflow benchmark grounded in live ClawHub marketplace signals and scored with deterministic checks plus structured judging.

Metric: Pass Rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.666.7
2GPT-5.463.8
3Claude Sonnet 4.661.9
4GLM-561.9
5MiniMax-M2.754.3
6Gemini 3.1 Pro (Preview)53.3
7MiMo-V2-Pro53.3
8DeepSeek V3.251.4
9Qwen 3.6 Plus50.5
10MiniMax-M2.550.5
11Qwen 3.5 397B A17B49.5

Interactive version: theaggregate.ai/benchmark?slug=claw-eval-live · How the rankings work · Data refreshed daily, snapshot 2026-07-22.