Claw-Eval: leaderboard

Agentic task-completion benchmark for complex multi-step scenarios, evaluating whether models can plan, act, and recover in realistic workflows.

Metric: Score (%). Source: claw-eval.github.io. 26 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.681.4
2MiMo-V2.5-Pro80.6
3Step 3.7 Flash80.5
4Claude Opus 4.680.4
5GPT-5.478.4
6MiMo-V2.578.4
7Kimi K2.677.9
8Gemini 3.1 Pro (Preview)77.3
9GLM-5.177.3
10DeepSeek V4 Pro77.2
11MiMo-V2-Pro77
12Muse Spark76.3
13Qwen 3.6 Plus75
14Kimi K2.574.9
15GLM-5-Turbo74.4

Interactive version: theaggregate.ai/benchmark?slug=claw-eval-hx1n6f · How It Works · Data refreshed daily, snapshot 2026-10-09.