LiveClawBench - Pass^3: leaderboard

Metric: Pass^3 (x100): share of the 134 cases where all three runs have reward above 0.8, OpenClaw-style personal-assistant tasks run in reproducible full-stack mock services (22 stateful web applications, browser, APIs and shell), three runs per case, rewards from case verifiers and a DeepSeek-V3.2 rubric judge for open-ended cases; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 17 models tracked.

Top models

#ModelScore
1GPT-5.553.7
2Kimi K2.7 Code51.5
3GLM-5.149.3
4MiniMax-M2.747
5Qwen 3.6 Plus46.3
6GLM-5.246.3
7Claude Opus 4.845.5
8MiMo-V2.5-Pro43.3
9Qwen 3.6 Flash43.3
10MiniMax-M342.5
11Kimi K2.639.1
12DeepSeek V4 Pro32.8
13DeepSeek V4 Flash29.9
14Qwen 3.6 27B29.9
15Qwen 3.5 Flash23.1

Interactive version: theaggregate.ai/benchmark?slug=liveclawbench-pass-pow-3 · How It Works · Data refreshed daily, snapshot 2026-10-07.