OptiVerse - Medium: leaderboard

Metric: Accuracy (%) on the 400 medium-level problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Flash53.25
2Gemini 3 Pro52.75
3GPT-5.2 (Thinking)50.75
4O347.25
5O4 Mini46.75
6Claude Sonnet 4.5 (Thinking)45.25
7DeepSeek V3.2 (Thinking)44.5
8Gemini 2.5 Pro43.75
9Gemini 2.5 Flash42.75
10GPT-OSS-120B39.25
11DeepSeek V3.2 (Non-reasoning)39.25
12Qwen 3 8B (Thinking)33
13Kimi K231.5
14Qwen 3 Coder 30B A3B Instruct19.25
15Ministral-3-8B-Instruct-251214.75

Interactive version: theaggregate.ai/benchmark?slug=optiverse-medium · How It Works · Data refreshed daily, snapshot 2026-10-07.