DynaSchedBench: leaderboard
Metric: Mean relative makespan gap (%): makespan above the best schedule observed for the same instance (all LLMs and 24 fixed priority dispatching rules), averaged over the 70-instance DynaSched-Subset of calibrated dynamic flexible job shop instances; LLM dispatches step by step with the L1 local observation, direct answer, two few-shot examples, temperature 0; 0 means matching the best schedule on every instance and there is no upper limit; lower is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 8B | 1.01 |
| 2 | Qwen 3 14B | 1.34 |
| 3 | Qwen 3 4B | 1.45 |
| 4 | Kimi K2 | 1.46 |
| 5 | Claude Haiku 4.5 | 1.52 |
| 6 | Qwen 3 32B | 1.65 |
| 7 | DeepSeek V3.2 | 1.68 |
| 8 | Gemini 2.5 Flash Lite | 1.76 |
| 9 | Grok 4 Fast | 1.88 |
| 10 | GPT-5 Mini | 1.93 |
Interactive version: theaggregate.ai/benchmark?slug=dynaschedbench · How It Works · Data refreshed daily, snapshot 2026-10-07.