GT Bench (Graph Theory) - Easy: leaderboard
Metric: Accuracy (%; Easy subset: connectivity and bipartiteness always, other tasks on sparse graphs or trees; zero-shot vanilla prompting at temperature 0.01, one run, over the 10,000 evaluation instances of 24 classical graph problems posed in four input representations (natural language, structured language, adjacency matrix, adjacency list); accuracy against the algorithmically verified answer). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 Mini (2025-01-31) | 98.66 |
| 2 | DeepSeek R1 | 96.32 |
| 3 | QwQ-32B | 85.62 |
| 4 | GPT-4o (2024-08-06) | 60.54 |
| 5 | Llama 3.3 70B Instruct | 58.19 |
| 6 | Phi-4 | 53.51 |
| 7 | GPT-4o Mini (2024-07-18) | 44.15 |
| 8 | Llama 3.1 8B Instruct | 38.13 |
Interactive version: theaggregate.ai/benchmark?slug=gt-bench-graph-theory-easy · How It Works · Data refreshed daily, snapshot 2026-09-26.