WildClawBench — leaderboard

WildClawBench evaluates model capability on agentic tasks from the linked upstream source with Overall Score as the primary reported metric.

Metric: Overall Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1Claude Opus 4.762.2
2GPT-5.558.2
3Nex N2 Pro53.5
4Claude Opus 4.651.6
5GPT-5.450.3
6GLM-5.148.2
7DeepSeek V4 Pro43.7
8GLM-542.6
9Gemini 3.1 Pro (Preview)40.8
10MiMo-V2-Pro40.2
11Qwen 3.5 397B A17B34.5
12DeepSeek V3.234
13GLM-5-Turbo33.9
14MiniMax-M2.733.8
15MiMo-V2-Flash30.8

Interactive version: theaggregate.ai/benchmark?slug=wildclawbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.