GraphARC - Output Graph Edge Count: leaderboard

Metric: Accuracy (%) of answers to questions about the number of edges of the output graph that the inferred transformation produces (the model never sees that output), GraphARC few-shot graph-transformation tasks (21 transformations; the model sees a few input-output graph pairs and a test input graph of 5 to 15 nodes encoded as an adjacency or incidence list), answer compared with the value computed from the true output graph; the OpenAI reasoning models ran one system prompt at medium reasoning effort, the other models four system-prompt variants; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-598
2O4 Mini98
3O3 Mini94
4O1 Mini76
5Qwen 3 14B67
6Qwen 3 8B61
7Qwen 3 32B55
8Qwen 3 4B54
9GPT-4.1 Nano46
10Qwen 3 1.7B27
11Llama 3.1 8B11
12DeepSeek R1 Distill Llama 8B11
13Llama 3 8B6
14Mistral-7B-v0.24
15OLMo 2 7B4

Interactive version: theaggregate.ai/benchmark?slug=grapharc-output-graph-edge-count · How It Works · Data refreshed daily, snapshot 2026-10-07.