OptiVerse - Easy: leaderboard

Metric: Accuracy (%) on the 300 easy-level problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1GPT-5.2 (Thinking)91
2Gemini 3 Pro89
3Gemini 3 Flash88.67
4O4 Mini87.67
5Gemini 2.5 Pro87
6O386.67
7DeepSeek V3.2 (Thinking)84.33
8Claude Sonnet 4.5 (Thinking)83.67
9Gemini 2.5 Flash82.33
10DeepSeek V3.2 (Non-reasoning)79.67
11GPT-OSS-120B78.67
12Qwen 3 8B (Thinking)73
13Kimi K271.33
14Qwen 3 Coder 30B A3B Instruct49.67
15Qwen 3 8B (Non-reasoning)42.67

Interactive version: theaggregate.ai/benchmark?slug=optiverse-easy · How It Works · Data refreshed daily, snapshot 2026-10-07.