RealClawBench - Code Fixing: leaderboard

Metric: Verifier pass rate (%) on the 63 code-fixing tasks of 281 released tasks reconstructed from real OpenClaw developer-agent sessions, run in the shared OpenClaw agent runtime (file read, write, edit, search and shell tools, 600-second timeout, temperature 1.0) and scored by case-specific deterministic Python verifiers; mean of three independent runs; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 14 models tracked.

Top models

#ModelScore
1MiMo-V2.5-Pro61.4
2DeepSeek V4 Pro60.3
3Claude Opus 4.758.7
4GPT-5.557.7
5Kimi K2.656.1
6Gemini 3.1 Pro (Preview)52.4
7GLM-5.151.9
8Claude Opus 4.650.3
9DeepSeek V4 Flash50.3
10Claude Sonnet 4.649.7
11MiniMax-M2.749.2
12GPT-OSS-120B47.6
13Qwen 3.6 Plus47.6
14Gemma 4 31B47.6

Interactive version: theaggregate.ai/benchmark?slug=realclawbench-code-fixing · How It Works · Data refreshed daily, snapshot 2026-09-29.