MM-OptBench - Multi-Period and System Planning: leaderboard
Metric: pass@1 (%) on the multi-period and system planning (120 instances) family, mean of five runs, solver-grounded: the model reads a text specification plus visual artifacts (tables, graphs, maps, schedules) and must return an optimization formulation and executable solver code whose returned objective matches the verified optimum, prompting only with no fine-tuning, retrieval or agentic repair; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 43.8 |
| 2 | Gemini 3.1 Pro (Preview) | 33.7 |
| 3 | Claude Sonnet 4.6 | 26.8 |
| 4 | GLM-4.5V | 10.2 |
Interactive version: theaggregate.ai/benchmark?slug=mm-optbench-multi-period-and-system-planning · How It Works · Data refreshed daily, snapshot 2026-10-07.