ClawProBench: leaderboard

OpenClaw agent benchmark measuring model performance on reasoning, planning, tool use, reliability, efficiency, and safety across repeated runs.

Metric: Overall Score (%). Source: suyoumo.github.io. Status: saturation imminent. 68 models tracked.

Top models

#ModelScore
1Kimi K381.6
2GLM-5.281.3
3MiniMax-M375.1
4Step 3.7 Flash72
5Ernie 5.170.7
6Qwen 3.5 397B A17B70.4
7Qwen 3.5 Plus70.1
8DeepSeek V4 Pro69.6
9GPT-5.5 (xHigh)69.3
10GLM-5.169
11MiMo-V2.5-Pro68.5
12Seed 2.0 Pro68.3
13GPT-5.4 (xHigh)68
14DeepSeek V3.267.6
15DeepSeek V4 Flash67.6

Interactive version: theaggregate.ai/benchmark?slug=clawprobench · How It Works · Data refreshed daily, snapshot 2026-09-05.