NPHardEval — leaderboard

Evaluates LLMs on 9 NP-hard computational problems (TSP, graph coloring, knapsack, etc.) spanning P, NP-complete, and NP-hard complexity classes.

Metric: Average Accuracy (%). Source: huggingface.co. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1GPT-4 Turbo37.74
2Claude 226.42
3GPT-3.5 Turbo25.86
4Claude Instant23.94
5Qwen-14B-Chat22.28
6Yi 34B (Chat)18.93
7Mistral 7B Instruct (v0.1)18.22
8Phi-29.21
9vicuna-13B-v1.37.56
10Phi-1.50.22
11mpt-30B Instruct0.02

Interactive version: theaggregate.ai/benchmark?slug=nphardeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.