OptiVerse: leaderboard

Metric: Accuracy (%) over all 1,000 OptiVerse problems; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Pro55.9
2Gemini 3 Flash55.5
3GPT-5.2 (Thinking)55.2
4O4 Mini51.2
5O351.1
6Claude Sonnet 4.5 (Thinking)49.7
7Gemini 2.5 Pro49.6
8DeepSeek V3.2 (Thinking)49.5
9Gemini 2.5 Flash47.4
10DeepSeek V3.2 (Non-reasoning)45.3
11GPT-OSS-120B44.9
12Qwen 3 8B (Thinking)39.9
13Kimi K238.9
14Qwen 3 Coder 30B A3B Instruct26.1
15Qwen 3 8B (Non-reasoning)20

Interactive version: theaggregate.ai/benchmark?slug=optiverse · How It Works · Data refreshed daily, snapshot 2026-10-07.