DynaSchedBench: leaderboard

Metric: Mean relative makespan gap (%): makespan above the best schedule observed for the same instance (all LLMs and 24 fixed priority dispatching rules), averaged over the 70-instance DynaSched-Subset of calibrated dynamic flexible job shop instances; LLM dispatches step by step with the L1 local observation, direct answer, two few-shot examples, temperature 0; 0 means matching the best schedule on every instance and there is no upper limit; lower is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.

Top models

#ModelScore
1Qwen 3 8B1.01
2Qwen 3 14B1.34
3Qwen 3 4B1.45
4Kimi K21.46
5Claude Haiku 4.51.52
6Qwen 3 32B1.65
7DeepSeek V3.21.68
8Gemini 2.5 Flash Lite1.76
9Grok 4 Fast1.88
10GPT-5 Mini1.93

Interactive version: theaggregate.ai/benchmark?slug=dynaschedbench · How It Works · Data refreshed daily, snapshot 2026-10-07.