TeamBench - Full Team: leaderboard

Metric: Pass rate (%) on the 90-task stratified TeamBench leaderboard subset (operating-system-enforced role separation in Docker sandboxes; a run passes when every check of the task's deterministic shell-script grader passes, and a run failing only the attestation check counts as a pass), temperature 0, with a Planner (reads the specification), an Executor (edits the workspace) and a Verifier (writes the final attestation), all the same model; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.737.8
2Gemini 3.1 Pro (Preview)28.9
3Claude Haiku 4.528.9
4GPT-5.4 Mini28.9
5Claude Sonnet 4.627.8
6GPT-5.427.8
7Gemini 3 Flash25.6
8Gemma 4 31B22.2
9Gemini 3.1 Flash Lite17.8

Interactive version: theaggregate.ai/benchmark?slug=teambench-full-team · How It Works · Data refreshed daily, snapshot 2026-10-07.