OptiVerse - Optimal Control: leaderboard

Metric: Accuracy (%) on the 78 optimal control problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1GPT-5.2 (Thinking)50
2Gemini 2.5 Flash48.72
3Gemini 3 Flash48.72
4Gemini 3 Pro42.31
5O4 Mini42.31
6O339.74
7Gemini 2.5 Pro38.46
8Claude Sonnet 4.5 (Thinking)34.62
9GPT-OSS-120B32.05
10DeepSeek V3.2 (Thinking)28.21
11Qwen 3 8B (Thinking)25.64
12Kimi K224.36
13DeepSeek V3.2 (Non-reasoning)21.79
14Qwen 3 Coder 30B A3B Instruct8.97
15Qwen 3 8B (Non-reasoning)6.41

Interactive version: theaggregate.ai/benchmark?slug=optiverse-optimal-control · How It Works · Data refreshed daily, snapshot 2026-10-07.