OR-Space - Revise: leaderboard

Metric: Pass@1 (%) on Revise-code: update the heuristic source to a revised business requirement, checked against the Gurobi oracle, 100 industrial optimization instances in a filesystem workspace (docs, data and an empty or heuristic src directory), Gurobi 12.0.1 track, API temperature at most 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)81
2Claude Opus 4.680
3GPT-5.479
4Claude Sonnet 4.578
5GPT-5 Mini77
6Claude Sonnet 4.5 (Thinking)76
7DeepSeek V4 Pro71
8GPT-5.171
9Gemini 2.5 Pro69
10Gemini 2.5 Flash66
11DeepSeek R1 052865
12DeepSeek V4 Flash63
13Qwen 3 Max60
14Qwen 3.5 27B48
15Gemini 3 Flash45

Interactive version: theaggregate.ai/benchmark?slug=or-space-revise · How It Works · Data refreshed daily, snapshot 2026-10-07.