OPT-BENCH-NP: leaderboard
Metric: Expert-gap closure (EG) at 20 optimization steps with history (OPT-Agent): per problem, (current metric minus initial solution metric) divided by (human expert heuristic metric minus initial metric), averaged over the 10 NP-hard problems (graph coloring, Hamiltonian cycle, knapsack, maximum clique, maximum set, meeting schedule, minimum cut, set cover, subset sum, TSP), times 100; 0 is the initial solution and 100 the expert heuristic, and a run may fall below the one or exceed the other (a signed, unbounded scale); higher is better. Source: arxiv.org. Saturation forecast: Around November 2026. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3.1 (Thinking) | 79 |
| 2 | Qwen 3 235B A22B (Thinking) | 76 |
| 3 | O3 Mini | 73 |
| 4 | Qwen 3 30B A3B | 69 |
| 5 | Claude 3.7 Sonnet (20250219) | 66 |
| 6 | Qwen 3 8B | 63 |
| 7 | Grok 3 | 62 |
| 8 | Qwen 3 32B | 60 |
| 9 | GPT-4.1 | 53 |
| 10 | Claude 3.5 Sonnet (20241022) | 52 |
| 11 | Qwen 2.5 32B Instruct | 52 |
| 12 | Qwen 2.5 72B Instruct | 47 |
| 13 | GPT-4o (2024-08-06) | 47 |
| 14 | Qwen 2.5 14B Instruct | 47 |
| 15 | Gemini 2.0 Flash | 45 |
Interactive version: theaggregate.ai/benchmark?slug=opt-bench-np · How It Works · Data refreshed daily, snapshot 2026-10-07.