OpenClawProBench — leaderboard

Agentic intelligence benchmark with 102 active scenarios testing planning, tool use, constraints, error recovery, synthesis, and safety. Deterministic grading across 3 trials per model.

Metric: Overall Score (%). Source: suyoumo.github.io. Status: saturation imminent. 68 models tracked.

Top models

#ModelScore
1Kimi K381.6
2GLM-5.281.3
3MiniMax-M375.1
4Step 3.7 Flash72
5Qwen 3.5 397B A17B70.4
6Qwen 3.5 Plus70.1
7DeepSeek V4 Pro69.6
8GPT-5.5 (xHigh)69.3
9GLM-5.169
10MiMo-V2.5-Pro68.5
11Seed 2.0 Pro68.3
12GPT-5.4 (xHigh)68
13DeepSeek V3.267.6
14DeepSeek V4 Flash67.6
15GLM-5-Turbo67.4

Interactive version: theaggregate.ai/benchmark?slug=openclawprobench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.