Claw-Anything - Execution Score: leaderboard
Metric: Average continuous execution score (0-1, the rubric soft score weighted toward final outcomes), over the 200 Claw-Anything tasks (simulated months of user activity, interdependent backend services, CLI Linux and GUI Android devices, reactive and proactive tasks), every model in the OpenHarness agent scaffold, three independent runs, rubric checks plus a Claude Sonnet 4.5 judge giving a pass label and a soft score; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 0.65 |
| 2 | Claude Opus 4.7 | 0.62 |
| 3 | Claude Sonnet 4.5 | 0.59 |
| 4 | GLM-5.1 | 0.59 |
| 5 | Qwen 3.6 27B | 0.58 |
| 6 | Kimi K2.6 | 0.57 |
| 7 | MiniMax-M2.7 | 0.52 |
| 8 | Qwen 3.5 27B | 0.5 |
Interactive version: theaggregate.ai/benchmark?slug=claw-anything-execution-score · How It Works · Data refreshed daily, snapshot 2026-10-07.