LiveClawBench - Easy: leaderboard

Metric: Average reward (0-100) over three runs on the 69 Easy cases (DeepSeek-V4-Pro anchor reward above 0.7), OpenClaw-style personal-assistant tasks run in reproducible full-stack mock services (22 stateful web applications, browser, APIs and shell), three runs per case, rewards from case verifiers and a DeepSeek-V3.2 rubric judge for open-ended cases; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScore
1Kimi K2.7 Code95.6
2Qwen 3.6 Plus95.3
3GLM-5.195.3
4MiniMax-M2.794.8
5Qwen 3.6 Flash94.3
6GPT-5.593.4
7Claude Opus 4.893.1
8MiniMax-M392.3
9Kimi K2.691.8
10GLM-5.291.6
11MiMo-V2.5-Pro90.7
12Qwen 3.6 27B82.3
13Qwen 3.5 Plus82
14DeepSeek V4 Flash81.9
15Qwen 3.5 Flash79.5

Interactive version: theaggregate.ai/benchmark?slug=liveclawbench-easy · How It Works · Data refreshed daily, snapshot 2026-10-07.