OptiVerse - Hard: leaderboard

Metric: Accuracy (%) on the 300 hard-level problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Pro27
2Gemini 3 Flash25.33
3GPT-5.2 (Thinking)25.33
4Claude Sonnet 4.5 (Thinking)21.67
5DeepSeek V3.2 (Thinking)21.33
6O320.67
7O4 Mini20.67
8Gemini 2.5 Pro20
9DeepSeek V3.2 (Non-reasoning)19
10Gemini 2.5 Flash18.67
11GPT-OSS-120B18.67
12Kimi K216.33
13Qwen 3 8B (Thinking)16
14Qwen 3 Coder 30B A3B Instruct11.67
15Qwen 3 8B (Non-reasoning)7.67

Interactive version: theaggregate.ai/benchmark?slug=optiverse-hard · How It Works · Data refreshed daily, snapshot 2026-10-07.