Claw-Anything - Consistency (Pass^3): leaderboard

Metric: Pass^3 (%): share of tasks solved in all three runs, over the 200 Claw-Anything tasks (simulated months of user activity, interdependent backend services, CLI Linux and GUI Android devices, reactive and proactive tasks), every model in the OpenHarness agent scaffold, three independent runs, rubric checks plus a Claude Sonnet 4.5 judge giving a pass label and a soft score; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 8 models tracked.

Top models

#ModelScore
1GPT-5.520
2GLM-5.117
3Claude Opus 4.713.5
4Claude Sonnet 4.512
5Kimi K2.66.5
6Qwen 3.6 27B6
7MiniMax-M2.73.5
8Qwen 3.5 27B2

Interactive version: theaggregate.ai/benchmark?slug=claw-anything-consistency-pass-pow-3 · How It Works · Data refreshed daily, snapshot 2026-10-07.