OptiVerse - Combinatorial Optimization: leaderboard

Metric: Accuracy (%) on the 238 combinatorial optimization problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Flash59.24
2Gemini 3 Pro57.14
3GPT-5.2 (Thinking)57.14
4O4 Mini54.2
5O353.78
6Claude Sonnet 4.5 (Thinking)52.52
7DeepSeek V3.2 (Thinking)52.52
8Gemini 2.5 Pro49.16
9DeepSeek V3.2 (Non-reasoning)47.48
10Gemini 2.5 Flash46.22
11Kimi K244.54
12Qwen 3 8B (Thinking)43.7
13Qwen 3 Coder 30B A3B Instruct31.09
14Qwen 3 8B (Non-reasoning)22.69
15InternLM3-8B-Instruct7.56

Interactive version: theaggregate.ai/benchmark?slug=optiverse-combinatorial-optimization · How It Works · Data refreshed daily, snapshot 2026-10-07.