LiveClawBench - Pass@3: leaderboard

Metric: Pass@3 (x100): share of the 134 cases where at least one of three runs has reward above 0.8, OpenClaw-style personal-assistant tasks run in reproducible full-stack mock services (22 stateful web applications, browser, APIs and shell), three runs per case, rewards from case verifiers and a DeepSeek-V3.2 rubric judge for open-ended cases; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1Kimi K2.7 Code70.9
2GLM-5.169.4
3MiniMax-M368.7
4GPT-5.567.9
5DeepSeek V4 Pro67.9
6MiniMax-M2.766.4
7GLM-5.266.4
8Kimi K2.665.4
9Qwen 3.6 Flash64.2
10Qwen 3.6 Plus63.4
11MiMo-V2.5-Pro63.4
12Claude Opus 4.861.9
13DeepSeek V4 Flash61.9
14Qwen 3.5 Plus59.7
15Qwen 3.6 27B58.2

Interactive version: theaggregate.ai/benchmark?slug=liveclawbench-pass-3 · How It Works · Data refreshed daily, snapshot 2026-10-07.