GT Bench (Graph Theory) - Hard: leaderboard

Metric: Accuracy (%; Hard subset: min-cost max-flow, cycle count, spanning-tree count and biconnected components always, other tasks on dense graphs; zero-shot vanilla prompting at temperature 0.01, one run, over the 10,000 evaluation instances of 24 classical graph problems posed in four input representations (natural language, structured language, adjacency matrix, adjacency list); accuracy against the algorithmically verified answer). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1O3 Mini (2025-01-31)75.09
2DeepSeek R164.1
3QwQ-32B38.46
4Phi-432.97
5Llama 3.3 70B Instruct30.4
6GPT-4o (2024-08-06)28.94
7GPT-4o Mini (2024-07-18)27.47
8Llama 3.1 8B Instruct19.05

Interactive version: theaggregate.ai/benchmark?slug=gt-bench-graph-theory-hard · How It Works · Data refreshed daily, snapshot 2026-09-26.