RealClawBench - Command Execution: leaderboard

Metric: Verifier pass rate (%) on the 24 command-execution tasks of 281 released tasks reconstructed from real OpenClaw developer-agent sessions, run in the shared OpenClaw agent runtime (file read, write, edit, search and shell tools, 600-second timeout, temperature 1.0) and scored by case-specific deterministic Python verifiers; mean of three independent runs; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 14 models tracked.

Top models

#ModelScore
1Kimi K2.666.7
2GLM-5.166.7
3GPT-5.565.3
4Claude Opus 4.765.3
5DeepSeek V4 Pro63.9
6MiMo-V2.5-Pro63.9
7DeepSeek V4 Flash62.5
8Gemini 3.1 Pro (Preview)59.7
9Claude Opus 4.658.3
10GPT-OSS-120B56.9
11MiniMax-M2.755.6
12Qwen 3.6 Plus54.2
13Claude Sonnet 4.652.8
14Gemma 4 31B52.8

Interactive version: theaggregate.ai/benchmark?slug=realclawbench-command-execution · How It Works · Data refreshed daily, snapshot 2026-09-29.