SCHEDBench (Solver Code): leaderboard

Metric: Feasibility rate (%; share of SCHEDBench natural-language scheduling instances whose model output satisfies all hard constraints, verified by OR-Tools CP-SAT; six canonical job-shop, project-scheduling, nurse-rostering and timetabling families; greedy decoding, zero-shot, no external solver; the model writes solver code executed in a sandbox rather than a schedule directly; full surface variation). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.4 (2026-03-05)39.6
2Qwen 3.5 397B A17B27.5
3Gemini 3.1 Flash Lite20.6
4Qwen 3.5 122B A10B19.4
5Qwen 3.5 27B14.6
6Claude Haiku 4.56
7Llama 4 Maverick1.3
8Llama 3.3 70B0.6

Interactive version: theaggregate.ai/benchmark?slug=schedbench-solver-code · How It Works · Data refreshed daily, snapshot 2026-09-26.