OptiVerse - Game Optimization: leaderboard

Metric: Accuracy (%) on the 51 game optimization problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Flash56.86
2Gemini 2.5 Flash54.9
3GPT-5.2 (Thinking)54.9
4Gemini 3 Pro50.98
5DeepSeek V3.2 (Thinking)49.02
6O4 Mini47.06
7Claude Sonnet 4.5 (Thinking)45.1
8O345.1
9Gemini 2.5 Pro43.14
10Qwen 3 8B (Thinking)37.25
11GPT-OSS-120B35.29
12DeepSeek V3.2 (Non-reasoning)35.29
13Kimi K225.49
14Qwen 3 8B (Non-reasoning)13.73
15Qwen 3 Coder 30B A3B Instruct9.8

Interactive version: theaggregate.ai/benchmark?slug=optiverse-game-optimization · How It Works · Data refreshed daily, snapshot 2026-10-07.