Pi-Bench - Completeness: leaderboard
Metric: Completeness (%): share of a task checklist items the final artifacts satisfy (deterministic or rubric graders), averaged over tasks; 100 multi-turn tasks over five professional personas in persistent workspaces, all models under the same agent scaffold adapted from nanobot with thinking enabled and default decoding, a GPT-5.4 simulated user and GPT-5.4 rubric grader, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | nanobot + Claude Opus 4.6 (Thinking) | 67.6 |
| 2 | nanobot + GPT-5.4 (Thinking) | 65.6 |
| 3 | nanobot + Qwen 3.6 Plus | 64.1 |
| 4 | nanobot + GLM-5.1 | 63.6 |
| 5 | nanobot + Kimi K2.5 | 61.6 |
| 6 | nanobot + Gemini 3.1 Pro (Preview) | 60 |
| 7 | nanobot + MiniMax-M2.7 | 60 |
| 8 | nanobot + DeepSeek V3.2 (Thinking) | 57.8 |
| 9 | nanobot + Seed 2.0 Pro | 52.1 |
Interactive version: theaggregate.ai/benchmark?slug=pi-bench-completeness · How It Works · Data refreshed daily, snapshot 2026-10-07.