OptiVerse - Stochastic Optimization: leaderboard

Metric: Accuracy (%) on the 120 stochastic optimization problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Pro56.67
2Gemini 3 Flash54.17
3Gemini 2.5 Pro50
4GPT-5.2 (Thinking)50
5Gemini 2.5 Flash48.33
6O4 Mini48.33
7Claude Sonnet 4.5 (Thinking)46.67
8O345.83
9DeepSeek V3.2 (Thinking)45
10DeepSeek V3.2 (Non-reasoning)45
11Qwen 3 8B (Thinking)41.67
12Kimi K238.33
13Qwen 3 Coder 30B A3B Instruct23.33
14Qwen 3 8B (Non-reasoning)20.83
15InternLM3-8B-Instruct5

Interactive version: theaggregate.ai/benchmark?slug=optiverse-stochastic-optimization · How It Works · Data refreshed daily, snapshot 2026-10-07.