ProactBench - Recovery: leaderboard

Metric: Pass rate (%) on the Recovery triggers (turns 8-10: grounded forward-looking value tied to an earlier detail after the user signals completion) over the 198 curated dialogues of ProactBench (about 624 rubric-scored trigger points: 201 Emergent, 232 Critical, 191 Recovery), each trigger judged Pass, Partial or Fail by a GPT-5.4 judge against a rubric written before the model answers; dialogue histories are fixed from a curation run and the model regenerates only at the triggers, sampled once at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 16 models tracked.

Top models

#ModelScore
1GPT-5.537.2
2Claude Opus 4.722.5
3DeepSeek V4 Flash18.3
4O4 Mini (2025-04-16)16.2
5MiMo-V2.5-Pro13.1
6Gemini 3.1 Pro (Preview)11.5
7Gemini 2.5 Pro (Medium)10.5
8GPT-4o (2024-11-20)8.9
9Kimi K2.67.4
10Qwen 2.5 7B Instruct6.8
11Gemini 2.5 Flash5.2
12Qwen 3 1.7B4.2
13Llama 4 Maverick2.1

Interactive version: theaggregate.ai/benchmark?slug=proactbench-recovery · How It Works · Data refreshed daily, snapshot 2026-10-07.