Claw-Eval-Live: leaderboard

Quarterly refreshed enterprise-workflow benchmark grounded in live ClawHub marketplace signals and scored with deterministic checks plus structured judging.

Metric: Pass Rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.666.7
2GPT-5.463.8
3Claude Sonnet 4.661.9
4GLM-561.9
5MiniMax-M2.754.3
6Gemini 3.1 Pro (Preview)53.3
7Kimi K2.553.3
8MiMo-V2-Pro53.3
9DeepSeek V3.251.4
10Qwen 3.6 Plus50.5
11MiniMax-M2.550.5
12Qwen 3.5 397B A17B49.5

Interactive version: theaggregate.ai/benchmark?slug=claw-eval-live · How It Works · Data refreshed daily, snapshot 2026-09-05.