ResearchClawBench — leaderboard

ResearchClawBench evaluates model capability on agentic tasks from the linked upstream source with Average Score as the primary reported metric.

Metric: Overall (self-reported). Source: benchmarklist.com. Status: saturation imminent. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.720.7
2Claude Opus 4.619.9
3Qwen 3.7 Max (Max)18.7
4GPT-5.418.4
5GLM-5.118.2
6Qwen 3.6 Plus18
7Gemini 3.5 Flash17.9
8DeepSeek V4 Pro17.1
9GPT-5.517
10MiMo-V2.516.9
11MiMo-V2-Pro15.3
12Qwen 3.5 397B A17B14.2
13Grok 4.113.5
14Gemini 3.1 Pro (Preview)13.3
15Grok 4.312.4

Interactive version: theaggregate.ai/benchmark?slug=researchclawbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.