AgentEscapeBench - Difficulty 10: leaderboard

Metric: Success rate (%; share of the 60 difficulty-10 instances (dependency DAGs of 10 tool and item nodes) whose deterministic final flag the agent submits; single run per instance, official APIs at temperature 0.7, multi-turn tool calling with incremental disclosure of hidden nodes). Source: arxiv.org. Saturation forecast: Around December 2026. 16 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)83.3
2GPT-5.483.3
3Claude Opus 4.683.3
4Kimi K2.581.7
5Gemini 3 Flash (Preview)66.7
6Seed 2.0 Pro63.3
7Qwen 3 235B A22B58.3
8MiniMax-M251.7
9GPT-548.3
10Qwen 3 Next 80B A3B35
11Qwen 3 32B16.7
12Qwen 3 14B11.7
13Qwen 3 8B0

Interactive version: theaggregate.ai/benchmark?slug=agentescapebench-difficulty-10 · How It Works · Data refreshed daily, snapshot 2026-09-26.