TeamBench - Team without Planner: leaderboard

Metric: Pass rate (%) on the 90-task stratified TeamBench leaderboard subset (operating-system-enforced role separation in Docker sandboxes; a run passes when every check of the task's deterministic shell-script grader passes, and a run failing only the attestation check counts as a pass), temperature 0, with an Executor and a Verifier of the same model, the Verifier holding the specification and the Executor seeing only the brief; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.735.6
2GPT-5.4 Mini25.6
3Gemma 4 31B24.4
4GPT-5.423.3
5Claude Haiku 4.518.9
6Gemini 3.1 Pro (Preview)16.7
7Gemini 3 Flash14.4
8GPT-OSS-20B12.2
9Claude Sonnet 4.610
10Gemini 3.1 Flash Lite8.9

Interactive version: theaggregate.ai/benchmark?slug=teambench-team-without-planner · How It Works · Data refreshed daily, snapshot 2026-10-07.