Pi-Bench - Completeness: leaderboard

Metric: Completeness (%): share of a task checklist items the final artifacts satisfy (deterministic or rubric graders), averaged over tasks; 100 multi-turn tasks over five professional personas in persistent workspaces, all models under the same agent scaffold adapted from nanobot with thinking enabled and default decoding, a GPT-5.4 simulated user and GPT-5.4 rubric grader, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.

Top models

#ModelScore
1nanobot + Claude Opus 4.6 (Thinking)67.6
2nanobot + GPT-5.4 (Thinking)65.6
3nanobot + Qwen 3.6 Plus64.1
4nanobot + GLM-5.163.6
5nanobot + Kimi K2.561.6
6nanobot + Gemini 3.1 Pro (Preview)60
7nanobot + MiniMax-M2.760
8nanobot + DeepSeek V3.2 (Thinking)57.8
9nanobot + Seed 2.0 Pro52.1

Interactive version: theaggregate.ai/benchmark?slug=pi-bench-completeness · How It Works · Data refreshed daily, snapshot 2026-10-07.