ProAgentBench - When to Assist: leaderboard
Metric: Accuracy (%) of the binary when-to-assist decision on ProAgentBench's held-out time-based split of real users' screen logs (positives are moments the user then queried an LLM, negatives contextually similar moments, about half each), from the preceding five minutes of OCR text, application metadata, user history and profile (text-only input), zero-shot prompt; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | DeepSeek V3.2 | 64.4 | #198 |
| 2 | Qwen 3 Max | 59.3 | #201 |
| 3 | Llama 3.1 8B Instruct | 57.3 | #1018 |
| 4 | GPT-4o Mini | 54.9 | #588 |
| 5 | Qwen 3 VL 8B Instruct | 51.7 | #401 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=proagentbench-when-to-assist · How It Works · Data refreshed daily, snapshot 2026-10-11.