RealClawBench - Data and Codebase Querying: leaderboard

Metric: Verifier pass rate (%) on the 88 data and codebase querying tasks of 281 released tasks reconstructed from real OpenClaw developer-agent sessions, run in the shared OpenClaw agent runtime (file read, write, edit, search and shell tools, 600-second timeout, temperature 1.0) and scored by case-specific deterministic Python verifiers; mean of three independent runs; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 14 models tracked.

Top models

#ModelScore
1Claude Opus 4.768.6
2GPT-5.564.4
3DeepSeek V4 Pro61.7
4Gemini 3.1 Pro (Preview)60.2
5MiMo-V2.5-Pro59.8
6Claude Sonnet 4.659.1
7GLM-5.158.7
8Kimi K2.658.3
9MiniMax-M2.757.2
10Qwen 3.6 Plus56.4
11Gemma 4 31B56.1
12DeepSeek V4 Flash55.7
13Claude Opus 4.651.5
14GPT-OSS-120B42.4

Interactive version: theaggregate.ai/benchmark?slug=realclawbench-data-and-codebase-querying · How It Works · Data refreshed daily, snapshot 2026-09-29.