OPT-BENCH-NP: leaderboard

Metric: Expert-gap closure (EG) at 20 optimization steps with history (OPT-Agent): per problem, (current metric minus initial solution metric) divided by (human expert heuristic metric minus initial metric), averaged over the 10 NP-hard problems (graph coloring, Hamiltonian cycle, knapsack, maximum clique, maximum set, meeting schedule, minimum cut, set cover, subset sum, TSP), times 100; 0 is the initial solution and 100 the expert heuristic, and a run may fall below the one or exceed the other (a signed, unbounded scale); higher is better. Source: arxiv.org. Saturation forecast: Around November 2026. 19 models tracked.

Top models

#ModelScore
1DeepSeek V3.1 (Thinking)79
2Qwen 3 235B A22B (Thinking)76
3O3 Mini73
4Qwen 3 30B A3B69
5Claude 3.7 Sonnet (20250219)66
6Qwen 3 8B63
7Grok 362
8Qwen 3 32B60
9GPT-4.153
10Claude 3.5 Sonnet (20241022)52
11Qwen 2.5 32B Instruct52
12Qwen 2.5 72B Instruct47
13GPT-4o (2024-08-06)47
14Qwen 2.5 14B Instruct47
15Gemini 2.0 Flash45

Interactive version: theaggregate.ai/benchmark?slug=opt-bench-np · How It Works · Data refreshed daily, snapshot 2026-10-07.