OptiVerse - Dynamic Optimization: leaderboard

Metric: Accuracy (%) on the 146 dynamic optimization problems of OptiVerse; the model writes and runs solver code (gurobi, pyomo, cvxpy, ortools and others) and a problem counts as solved only when every required variable and the objective match the ground truth within 0.1% relative error, extracted and checked by a DeepSeek-V3.2-Chat judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1GPT-5.2 (Thinking)57.53
2Gemini 3 Pro56.85
3O356.85
4Gemini 3 Flash55.48
5O4 Mini55.48
6DeepSeek V3.2 (Thinking)49.32
7Gemini 2.5 Flash48.63
8Gemini 2.5 Pro47.95
9Claude Sonnet 4.5 (Thinking)47.95
10DeepSeek V3.2 (Non-reasoning)43.15
11Kimi K239.73
12Qwen 3 8B (Thinking)39.73
13Qwen 3 Coder 30B A3B Instruct23.29
14Qwen 3 8B (Non-reasoning)14.38
15InternLM3-8B-Instruct8.22

Interactive version: theaggregate.ai/benchmark?slug=optiverse-dynamic-optimization · How It Works · Data refreshed daily, snapshot 2026-10-07.