ProAgentBench - When to Assist: leaderboard

Metric: Accuracy (%) of the binary when-to-assist decision on ProAgentBench's held-out time-based split of real users' screen logs (positives are moments the user then queried an LLM, negatives contextually similar moments, about half each), from the preceding five minutes of OCR text, application metadata, user history and profile (text-only input), zero-shot prompt; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek V3.264.4#198
2Qwen 3 Max59.3#201
3Llama 3.1 8B Instruct57.3#1018
4GPT-4o Mini54.9#588
5Qwen 3 VL 8B Instruct51.7#401

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=proagentbench-when-to-assist · How It Works · Data refreshed daily, snapshot 2026-10-11.