Pi-Bench - Proactivity: leaderboard

Metric: Proactivity (%): share of a task hidden intents the agent resolves without the user stating them, by acting on them or by asking a targeted clarification, averaged over tasks; 100 multi-turn tasks over five professional personas in persistent workspaces, all models under the same agent scaffold adapted from nanobot with thinking enabled and default decoding, a GPT-5.4 simulated user and GPT-5.4 rubric grader, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.

Top models

#ModelScore
1nanobot + GPT-5.4 (Thinking)67
2nanobot + Claude Opus 4.6 (Thinking)65.5
3nanobot + Qwen 3.6 Plus64
4nanobot + Seed 2.0 Pro58.4
5nanobot + GLM-5.158.4
6nanobot + Gemini 3.1 Pro (Preview)57.1
7nanobot + MiniMax-M2.755.6
8nanobot + DeepSeek V3.2 (Thinking)53.3
9nanobot + Kimi K2.543.1

Interactive version: theaggregate.ai/benchmark?slug=pi-bench-proactivity · How It Works · Data refreshed daily, snapshot 2026-10-07.