AgentCE-Bench - Travel Itinerary: leaderboard
Metric: Task score (%) on AgentCE-Bench travel domain (54 instances): the agent fills the hidden slots of a 5x7 planning grid with tool calls on static JSON (attribute filter queries, slot and global constraint checkers), each hidden slot having 25 candidates including decoys that violate only the global constraints; every cell is a whole count of instances over the 54 per domain (6 hidden-slot counts by 9 decoy budgets); at most 600 steps; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 397B A17B | 77.8 |
| 2 | Qwen 3.5 27B | 74.1 |
| 3 | Qwen 3.5 122B A10B | 74.1 |
| 4 | Qwen 3.5 35B A3B | 68.5 |
| 5 | MiniMax-M2.5 | 61.1 |
| 6 | MiniMax-M2.1 | 53.7 |
| 7 | Qwen 3.5 9B | 46.3 |
| 8 | MiniMax-M2 | 44.4 |
| 9 | GLM-4.7 FP8 | 40.7 |
| 10 | Qwen 3.5 4B | 33.3 |
| 11 | Qwen 3.5 2B | 1.9 |
| 12 | Qwen 3.5 0.8B | 0 |
Interactive version: theaggregate.ai/benchmark?slug=agentce-bench-travel-itinerary · How It Works · Data refreshed daily, snapshot 2026-10-07.