E-Bench-Code: leaderboard
Metric: Avg@3 (%; per-trial success averaged over tasks and three trials; domain MCP tools plus an exec_code tool) on 323 synthetic state-changing tasks in three product environments (Honor of Kings, QQ Music, Tencent Meeting), graded by exact database-state diff; three trials per task, each model at its highest available thinking effort in a shared MCP harness. Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 (Max) | 81.11 |
| 2 | Kimi K3 (Max) | 77.61 |
| 3 | GPT-5.5 (xHigh) | 77.19 |
| 4 | Grok 4.5 (High) | 69.24 |
| 5 | Hy3 | 64.4 |
| 6 | Gemini 3.5 Flash (High) | 62.54 |
| 7 | Qwen 3.7 Max | 61.92 |
| 8 | GLM-5.2 (Max) | 60.99 |
| 9 | Seed 2.1 Pro | 53.04 |
| 10 | DeepSeek V4 Pro (Max) | 47.68 |
| 11 | MiniMax-M3 | 46.85 |
Interactive version: theaggregate.ai/benchmark?slug=e-bench-code · How It Works · Data refreshed daily, snapshot 2026-09-29.