WorkBench 2026 - Side Effects: leaderboard

Metric: Side-effect rate (%): share of tasks in which the agent took a wrong, harmful action (for example emailing the wrong person) on the corrected 2026 release of WorkBench: 690 templated workplace tasks over a sandbox of five databases (calendar, email, website analytics, CRM, project board) and 26 tools, graded by comparing the final sandbox state with the ground truth; one native tool-calling agent loop for every model, up to 20 steps, temperature 0 where allowed; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 24 models tracked.

Top models

#ModelScore
1Claude Fable 51.9
2Claude Opus 4.82.5
3Gemini 3.1 Pro (Preview)2.9
4Gemini 3.5 Flash3
5GPT-5.53.9
6Kimi K2.66.8
7Claude Sonnet 4.68.8
8DeepSeek V4 Pro12.8
9GPT-513
10GPT-4o15.1
11Claude Haiku 4.516.7
12GPT-5.416.8
13GLM-4.617
14O317.5
15GPT-5.118.1

Interactive version: theaggregate.ai/benchmark?slug=workbench-2026-side-effects · How It Works · Data refreshed daily, snapshot 2026-09-29.