MM-OptBench - Combinatorial and Logical Models: leaderboard

Metric: pass@1 (%) on the pure combinatorial and logical models (120 instances) family, mean of five runs, solver-grounded: the model reads a text specification plus visual artifacts (tables, graphs, maps, schedules) and must return an optimization formulation and executable solver code whose returned objective matches the verified optimum, prompting only with no fine-tuning, retrieval or agentic repair; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScore
1GPT-5.482.3
2Gemini 3.1 Pro (Preview)49.2
3Claude Sonnet 4.633.8
4GLM-4.5V23.8

Interactive version: theaggregate.ai/benchmark?slug=mm-optbench-combinatorial-and-logical-models · How It Works · Data refreshed daily, snapshot 2026-10-07.