ProAgentBench (CoT) - When to Assist: leaderboard

Metric: Accuracy (%) of the binary when-to-assist decision on ProAgentBench's held-out time-based split of real users' screen logs (positives are moments the user then queried an LLM, negatives contextually similar moments, about half each), from the preceding five minutes of OCR text, application metadata, user history and profile (text-only input), task-specific chain-of-thought prompt; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 6 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek V3.261.1#198
2Qwen 3 Max59.8#201
3GPT-4o Mini55.7#588
4Llama 3.1 8B Instruct50.8#1018
5Qwen 3 VL 8B Instruct41#401

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=proagentbench-cot-when-to-assist · How It Works · Data refreshed daily, snapshot 2026-10-11.