ProactBench - Weighted Score: leaderboard

Metric: Weighted score (%): Pass counts 1, Partial 0.5 and Fail 0, averaged over all trigger points over the 198 curated dialogues of ProactBench (about 624 rubric-scored trigger points: 201 Emergent, 232 Critical, 191 Recovery), each trigger judged Pass, Partial or Fail by a GPT-5.4 judge against a rubric written before the model answers; dialogue histories are fixed from a curation run and the model regenerates only at the triggers, sampled once at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScore
1GPT-5.569.8
2Claude Opus 4.760.6
3Gemini 3.1 Pro (Preview)58.3
4Kimi K2.654.3
5DeepSeek V4 Flash53.5
6MiMo-V2.5-Pro52.8
7Gemini 2.5 Pro (Medium)50.3
8O4 Mini (2025-04-16)47.4
9Gemini 2.5 Flash44.2
10GPT-4o (2024-11-20)31.3
11Llama 4 Maverick22.5
12Qwen 3 1.7B22
13Qwen 2.5 7B Instruct21.6

Interactive version: theaggregate.ai/benchmark?slug=proactbench-weighted-score · How It Works · Data refreshed daily, snapshot 2026-10-07.