E-Bench - Pass^3: leaderboard

Metric: Pass^3 (%; tasks solved in all three trials; domain MCP tools only) on 323 synthetic state-changing tasks in three product environments (Honor of Kings, QQ Music, Tencent Meeting), graded by exact database-state diff; three trials per task, each model at its highest available thinking effort in a shared MCP harness. Source: arxiv.org. Saturation forecast: Around February 2027. 11 models tracked.

Top models

#ModelScore
1Kimi K3 (Max)58.82
2GPT-5.5 (xHigh)57.59
3Grok 4.5 (High)52.32
4Claude Opus 4.8 (Max)50.81
5GLM-5.2 (Max)30.96
6Qwen 3.7 Max30.34
7Seed 2.1 Pro28.79
8Hy326.63
9Gemini 3.5 Flash (High)21.98
10MiniMax-M320.12
11DeepSeek V4 Pro (Max)17.34

Interactive version: theaggregate.ai/benchmark?slug=e-bench-pass-pow-3 · How It Works · Data refreshed daily, snapshot 2026-09-29.