ProactBench - Critical: leaderboard

Metric: Pass rate (%) on the Critical triggers (turns 4-7: synthesising two or more disclosed anchors into a new conclusion) over the 198 curated dialogues of ProactBench (about 624 rubric-scored trigger points: 201 Emergent, 232 Critical, 191 Recovery), each trigger judged Pass, Partial or Fail by a GPT-5.4 judge against a rubric written before the model answers; dialogue histories are fixed from a curation run and the model regenerates only at the triggers, sampled once at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScore
1GPT-5.575.4
2Gemini 3.1 Pro (Preview)65.5
3Kimi K2.664.4
4Claude Opus 4.758.6
5MiMo-V2.5-Pro53.2
6DeepSeek V4 Flash50.6
7Gemini 2.5 Pro (Medium)48.3
8O4 Mini (2025-04-16)44.8
9Gemini 2.5 Flash37.9
10GPT-4o (2024-11-20)18.5
11Llama 4 Maverick14.2
12Qwen 2.5 7B Instruct6.5
13Qwen 3 1.7B6.5

Interactive version: theaggregate.ai/benchmark?slug=proactbench-critical · How It Works · Data refreshed daily, snapshot 2026-10-07.