RealClawBench - Project Building: leaderboard

Metric: Verifier pass rate (%) on the 49 project-building tasks of 281 released tasks reconstructed from real OpenClaw developer-agent sessions, run in the shared OpenClaw agent runtime (file read, write, edit, search and shell tools, 600-second timeout, temperature 1.0) and scored by case-specific deterministic Python verifiers; mean of three independent runs; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Claude Opus 4.770.1
2GPT-5.569.4
3MiMo-V2.5-Pro63.9
4DeepSeek V4 Pro61.9
5GLM-5.161.9
6Gemini 3.1 Pro (Preview)60.5
7Qwen 3.6 Plus59.2
8Claude Opus 4.658.5
9Claude Sonnet 4.655.8
10DeepSeek V4 Flash55.8
11Kimi K2.652.4
12MiniMax-M2.752.4
13Gemma 4 31B44.9
14GPT-OSS-120B38.1

Interactive version: theaggregate.ai/benchmark?slug=realclawbench-project-building · How It Works · Data refreshed daily, snapshot 2026-09-29.