Do Agents Know What They Can't Do? Evaluating — leaderboard

Metric: Avg. (self-reported). Source: benchmarklist.com. 9 models tracked.

Top models

#ModelScore
1GPT-5.561.2
2GPT-OSS-120B60.7
3Qwen 3.5 397B A17B58.2
4Qwen 3.5 27B56.7
5DeepSeek V4 Pro54.8
6DeepSeek V4 Flash52.9
7Qwen 3.5 122B A10B52.3
8Qwen 3.5 35B A3B52.3
9Qwen 3.5 9B47.2

Interactive version: theaggregate.ai/benchmark?slug=do-agents-know-what-they-can-t-do-evaluating · How the rankings work · Data refreshed daily, snapshot 2026-07-22.