WorkBench 2026: leaderboard

Metric: Successful task completion (%) on the corrected 2026 release of WorkBench: 690 templated workplace tasks over a sandbox of five databases (calendar, email, website analytics, CRM, project board) and 26 tools, graded by comparing the final sandbox state with the ground truth; one native tool-calling agent loop for every model, up to 20 steps, temperature 0 where allowed; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 24 models tracked.

Top models

#ModelScore
1Claude Fable 597.7
2Claude Opus 4.896.2
3Gemini 3.1 Pro (Preview)95.5
4GPT-5.594.9
5Gemini 3.5 Flash92
6Claude Sonnet 4.688.3
7Kimi K2.688
8GPT-584.3
9DeepSeek V4 Pro84.2
10GPT-5.478.6
11O377.4
12GLM-4.677.4
13Claude Haiku 4.574.8
14GPT-4.173.5
15GPT-5.270.4

Interactive version: theaggregate.ai/benchmark?slug=workbench-2026 · How It Works · Data refreshed daily, snapshot 2026-09-29.