Claw-Anything - CLI and GUI Tasks: leaderboard

Metric: Pass@1 (%) on the 50 tasks that need both the CLI and the GUI device, over the 200 Claw-Anything tasks (simulated months of user activity, interdependent backend services, CLI Linux and GUI Android devices, reactive and proactive tasks), every model in the OpenHarness agent scaffold, three independent runs, rubric checks plus a Claude Sonnet 4.5 judge giving a pass label and a soft score; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 8 models tracked.

Top models

#ModelScore
1Qwen 3.6 27B18
2GPT-5.517.3
3GLM-5.116
4MiniMax-M2.711.3
5Kimi K2.69.3
6Qwen 3.5 27B9.3
7Claude Opus 4.77.3
8Claude Sonnet 4.56

Interactive version: theaggregate.ai/benchmark?slug=claw-anything-cli-and-gui-tasks · How It Works · Data refreshed daily, snapshot 2026-10-07.