ProAgentBench (CoT) - When to Assist: leaderboard
Metric: Accuracy (%) of the binary when-to-assist decision on ProAgentBench's held-out time-based split of real users' screen logs (positives are moments the user then queried an LLM, negatives contextually similar moments, about half each), from the preceding five minutes of OCR text, application metadata, user history and profile (text-only input), task-specific chain-of-thought prompt; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | DeepSeek V3.2 | 61.1 | #198 |
| 2 | Qwen 3 Max | 59.8 | #201 |
| 3 | GPT-4o Mini | 55.7 | #588 |
| 4 | Llama 3.1 8B Instruct | 50.8 | #1018 |
| 5 | Qwen 3 VL 8B Instruct | 41 | #401 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=proagentbench-cot-when-to-assist · How It Works · Data refreshed daily, snapshot 2026-10-11.