ClawBench (OpenClaw): leaderboard

Metric: Success rate (%) over the 153 ClawBench tasks on 144 live platforms (agents in the OpenClaw browser harness on live production websites, the final state-changing request intercepted before it reaches the server, each run judged pass or fail by an Agent-as-Judge (Claude Sonnet 4.6) against a human reference trajectory); higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (High)33.3
2GLM-524.2
3Gemini 3 Flash19
4Claude Haiku 4.518.3
5Gemini 3.1 Pro (Preview)9.8
6GPT-5.4 (High)6.5
7Gemini 3.1 Flash Lite3.3

Interactive version: theaggregate.ai/benchmark?slug=clawbench-openclaw · How It Works · Data refreshed daily, snapshot 2026-10-07.