WildClawBench: leaderboard

60 hand-built agent tasks run inside the OpenClaw assistant with real shells, files, browsers, email and calendar, spanning productivity, coding, social, search, creative and safety; InternLM, 2026.

Metric: Overall Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 66 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol67.2
2Seed 2.1 Turbo62.8
3Claude Opus 4.762.2
4GPT-5.558.2
5Kimi K354.5
6Nex N2 Pro53.5
7Claude Opus 4.651.6
8GPT-5.450.3
9GLM-5.148.2
10Muse Glimmer 30B47.6
11Kimi K2.7 Code46.9
12DeepSeek V4 Pro43.7
13GLM-542.6
14Gemini 3.1 Pro (Preview)40.8
15MiMo-V2-Pro40.2

Interactive version: theaggregate.ai/benchmark?slug=wildclawbench · How It Works · Data refreshed daily, snapshot 2026-09-05.