ClawBench (OpenClaw) - Daily: leaderboard

Metric: Success rate (%) on the 52 daily-life tasks of the 153-task ClawBench paper evaluation (agents in the OpenClaw browser harness on live production websites, the final state-changing request intercepted before it reaches the server, each run judged pass or fail by an Agent-as-Judge (Claude Sonnet 4.6) against a human reference trajectory); higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (High)44.2
2GLM-530.8
3Claude Haiku 4.515.4
4Gemini 3 Flash15.4
5GPT-5.4 (High)9.6
6Gemini 3.1 Pro (Preview)5.8
7Gemini 3.1 Flash Lite1.9

Interactive version: theaggregate.ai/benchmark?slug=clawbench-openclaw-daily · How It Works · Data refreshed daily, snapshot 2026-10-07.