RealClawBench - File Creation: leaderboard

Metric: Verifier pass rate (%) on the 57 file-creation tasks of 281 released tasks reconstructed from real OpenClaw developer-agent sessions, run in the shared OpenClaw agent runtime (file read, write, edit, search and shell tools, 600-second timeout, temperature 1.0) and scored by case-specific deterministic Python verifiers; mean of three independent runs; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1GPT-5.570.2
2Claude Opus 4.766.1
3Qwen 3.6 Plus61.4
4DeepSeek V4 Flash60.2
5Kimi K2.659.1
6Claude Sonnet 4.656.1
7Gemini 3.1 Pro (Preview)55
8MiMo-V2.5-Pro54.4
9GLM-5.153.8
10MiniMax-M2.753.8
11DeepSeek V4 Pro52.6
12Claude Opus 4.646.8
13GPT-OSS-120B35.7
14Gemma 4 31B25.1

Interactive version: theaggregate.ai/benchmark?slug=realclawbench-file-creation · How It Works · Data refreshed daily, snapshot 2026-09-29.