Agent-Diff - Pass Rate: leaderboard

Metric: Pass rate (%) on Agent-Diff's 224 enterprise API tasks (Box, Google Calendar, Linear, Slack replicas) in the no-documentation setting, code-executing agent, three trials per task, graded by state-diff assertions with any unexpected side effect zeroing the task; a task passes when it is clean and every assertion holds; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek V3.276#198
2Devstral 274#394
3Gemini 3 Flash (Preview)67#78
4Kimi K2 090564#282
5GPT-OSS-120B60#330
6Grok 4.1 Fast52#208
7Claude Haiku 4.550#271
8Llama 4 Scout29#646

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=agent-diff-pass-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.