AgentCE-Bench: leaderboard

Metric: Task score (%) on AgentCE-Bench over all 324 instances, the mean of the six domains: the agent fills the hidden slots of a 5x7 planning grid with tool calls on static JSON (attribute filter queries, slot and global constraint checkers), each hidden slot having 25 candidates including decoys that violate only the global constraints; every cell is a whole count of instances over the 54 per domain (6 hidden-slot counts by 9 decoy budgets); at most 600 steps; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B84.9
2Qwen 3.5 122B A10B75
3Qwen 3.5 27B74.7
4Qwen 3.5 35B A3B69.1
5MiniMax-M2.558.6
6MiniMax-M2.157.1
7GLM-4.7 FP846.9
8Qwen 3.5 9B46
9MiniMax-M244.1
10Qwen 3.5 4B36.1
11Qwen 3.5 2B3.4
12Qwen 3.5 0.8B0.6

Interactive version: theaggregate.ai/benchmark?slug=agentce-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.