Pi-Bench - Proactivity: leaderboard
Metric: Proactivity (%): share of a task hidden intents the agent resolves without the user stating them, by acting on them or by asking a targeted clarification, averaged over tasks; 100 multi-turn tasks over five professional personas in persistent workspaces, all models under the same agent scaffold adapted from nanobot with thinking enabled and default decoding, a GPT-5.4 simulated user and GPT-5.4 rubric grader, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | nanobot + GPT-5.4 (Thinking) | 67 |
| 2 | nanobot + Claude Opus 4.6 (Thinking) | 65.5 |
| 3 | nanobot + Qwen 3.6 Plus | 64 |
| 4 | nanobot + Seed 2.0 Pro | 58.4 |
| 5 | nanobot + GLM-5.1 | 58.4 |
| 6 | nanobot + Gemini 3.1 Pro (Preview) | 57.1 |
| 7 | nanobot + MiniMax-M2.7 | 55.6 |
| 8 | nanobot + DeepSeek V3.2 (Thinking) | 53.3 |
| 9 | nanobot + Kimi K2.5 | 43.1 |
Interactive version: theaggregate.ai/benchmark?slug=pi-bench-proactivity · How It Works · Data refreshed daily, snapshot 2026-10-07.