ProactBench — leaderboard

Metric: Overall Pass Rate (self-reported). Source: benchmarklist.com. 16 models tracked.

Top models

#ModelScore
1GPT-5.561.5
2Qwen 3.5 397B A17B52.5
3Claude Opus 4.751.9
4Gemini 3.1 Pro (Preview)50.4
5Qwen 3.5 9B48
6MiMo-V2.5-Pro44.9
7DeepSeek V4 Flash44.3
8Gemini 2.5 Pro42.5
9O4 Mini38.6
10Gemini 2.5 Flash35.3
11GPT-4o20.5
12Llama 4 Maverick13.1
13Qwen 2.5 7B Instruct11.5

Interactive version: theaggregate.ai/benchmark?slug=proactbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.