OpenClawProBench: leaderboard

Agentic intelligence benchmark with 102 active scenarios testing planning, tool use, constraints, error recovery, synthesis, and safety. Deterministic grading across 3 trials per model.

Metric: Overall Score (%). Source: suyoumo.github.io. Status: years away from saturation. 68 models tracked.

Top models

#ModelScore
1Kimi K381.6
2GLM-5.281.3
3MiniMax-M375.1
4Step 3.7 Flash72
5Ernie 5.170.7
6Qwen 3.5 397B A17B70.4
7Qwen 3.5 Plus70.1
8DeepSeek V4 Pro69.6
9GPT-5.5 (xHigh)69.3
10GLM-5.169
11MiMo-V2.5-Pro68.5
12Seed 2.0 Pro68.3
13GPT-5.4 (xHigh)68
14DeepSeek V3.267.6
15DeepSeek V4 Flash67.6

Interactive version: theaggregate.ai/benchmark?slug=openclawprobench · How It Works · Data refreshed daily, snapshot 2026-09-05.