MM-OptBench (Pass@4): leaderboard

Metric: pass@4 (%) over the 780 instances, success in at least one of four sampled runs, solver-grounded: the model reads a text specification plus visual artifacts (tables, graphs, maps, schedules) and must return an optimization formulation and executable solver code whose returned objective matches the verified optimum, prompting only with no fine-tuning, retrieval or agentic repair; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1GPT-5.468.3
2Gemini 3.1 Pro (Preview)64.4
3Claude Sonnet 4.631.3
4GLM-4.5V29.1

Interactive version: theaggregate.ai/benchmark?slug=mm-optbench-pass-4 · How It Works · Data refreshed daily, snapshot 2026-10-07.