ClawProBench — leaderboard

OpenClaw agent benchmark measuring model performance on reasoning, planning, tool use, reliability, efficiency, and safety across repeated runs.

Metric: Final Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 57 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)67.9
2DeepSeek V4 Pro64.38
3Qwen 3.5 Plus64.19
4Qwen 3.5 397B A17B64.18
5MiMo-V2.5-Pro63.3
6GLM-5.162.93
7GLM-5-Turbo61.92
8DeepSeek V4 Flash61.47
9Seed 2.0 Pro61.07
10Claude Sonnet 4.660.5
11Seed 2.0 Lite60.4
12MiMo-V2.560.39
13Qwen 3.6 Plus60.2
14DeepSeek V3.260.13
15Kimi K2.659.31

Interactive version: theaggregate.ai/benchmark?slug=clawprobench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.