GT Bench (Graph Theory) - Easy: leaderboard

Metric: Accuracy (%; Easy subset: connectivity and bipartiteness always, other tasks on sparse graphs or trees; zero-shot vanilla prompting at temperature 0.01, one run, over the 10,000 evaluation instances of 24 classical graph problems posed in four input representations (natural language, structured language, adjacency matrix, adjacency list); accuracy against the algorithmically verified answer). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1O3 Mini (2025-01-31)98.66
2DeepSeek R196.32
3QwQ-32B85.62
4GPT-4o (2024-08-06)60.54
5Llama 3.3 70B Instruct58.19
6Phi-453.51
7GPT-4o Mini (2024-07-18)44.15
8Llama 3.1 8B Instruct38.13

Interactive version: theaggregate.ai/benchmark?slug=gt-bench-graph-theory-easy · How It Works · Data refreshed daily, snapshot 2026-09-26.