OptiVerse - Mathematical Programming: leaderboard

Metric: Accuracy (%) on the 367 mathematical programming problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Pro58.04
2GPT-5.2 (Thinking)55.86
3Gemini 3 Flash54.77
4Gemini 2.5 Pro53.68
5DeepSeek V3.2 (Thinking)53.68
6Claude Sonnet 4.5 (Thinking)53.41
7O352.04
8DeepSeek V3.2 (Non-reasoning)51.23
9O4 Mini50.95
10Gemini 2.5 Flash46.05
11Qwen 3 8B (Thinking)40.33
12Kimi K240.05
13Qwen 3 Coder 30B A3B Instruct30.79
14Qwen 3 8B (Non-reasoning)23.98
15InternLM3-8B-Instruct11.44

Interactive version: theaggregate.ai/benchmark?slug=optiverse-mathematical-programming · How It Works · Data refreshed daily, snapshot 2026-10-07.