AgentCE-Bench - Workforce Scheduling: leaderboard

Metric: Task score (%) on AgentCE-Bench workforce domain (54 instances): the agent fills the hidden slots of a 5x7 planning grid with tool calls on static JSON (attribute filter queries, slot and global constraint checkers), each hidden slot having 25 candidates including decoys that violate only the global constraints; every cell is a whole count of instances over the 54 per domain (6 hidden-slot counts by 9 decoy budgets); at most 600 steps; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B90.7
2Qwen 3.5 122B A10B75.9
3Qwen 3.5 27B68.5
4MiniMax-M2.164.8
5Qwen 3.5 35B A3B61.1
6MiniMax-M2.561.1
7GLM-4.7 FP850
8Qwen 3.5 9B46.3
9MiniMax-M244.4
10Qwen 3.5 4B25.9
11Qwen 3.5 2B3.7
12Qwen 3.5 0.8B0

Interactive version: theaggregate.ai/benchmark?slug=agentce-bench-workforce-scheduling · How It Works · Data refreshed daily, snapshot 2026-10-07.