OR-Space - Explain: leaderboard

Metric: Explain rubric score (0-100): exact coverage, reasoning, grounding and answer quality of short solver-grounded reports minus a hallucination penalty, checklist items scored programmatically or by a judge model, 100 industrial optimization instances in a filesystem workspace (docs, data and an empty or heuristic src directory), Gurobi 12.0.1 track, API temperature at most 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1GPT-5.486.52
2GPT-5 Mini82.94
3Claude Opus 4.680.2
4Claude Sonnet 4.5 (Thinking)79.63
5Qwen 3.5 27B77.88
6DeepSeek V4 Flash77.77
7Claude Sonnet 4.577.58
8GPT-5.176.16
9Qwen 3 Max74.66
10Gemini 3.1 Pro (Preview)73
11DeepSeek V4 Pro68.87
12Qwen 3 32B (Non-reasoning)61.32
13Qwen 3 32B61.28
14Qwen 3 14B (Non-reasoning)61.03
15GPT-4o59.36

Interactive version: theaggregate.ai/benchmark?slug=or-space-explain · How It Works · Data refreshed daily, snapshot 2026-10-07.