GT Bench (Graph Theory) - Hard: leaderboard
Metric: Accuracy (%; Hard subset: min-cost max-flow, cycle count, spanning-tree count and biconnected components always, other tasks on dense graphs; zero-shot vanilla prompting at temperature 0.01, one run, over the 10,000 evaluation instances of 24 classical graph problems posed in four input representations (natural language, structured language, adjacency matrix, adjacency list); accuracy against the algorithmically verified answer). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 Mini (2025-01-31) | 75.09 |
| 2 | DeepSeek R1 | 64.1 |
| 3 | QwQ-32B | 38.46 |
| 4 | Phi-4 | 32.97 |
| 5 | Llama 3.3 70B Instruct | 30.4 |
| 6 | GPT-4o (2024-08-06) | 28.94 |
| 7 | GPT-4o Mini (2024-07-18) | 27.47 |
| 8 | Llama 3.1 8B Instruct | 19.05 |
Interactive version: theaggregate.ai/benchmark?slug=gt-bench-graph-theory-hard · How It Works · Data refreshed daily, snapshot 2026-09-26.