GraphARC - Output Graph Component Count: leaderboard

Metric: Accuracy (%) of answers to questions about the number of connected components of the output graph that the inferred transformation produces (the model never sees that output), GraphARC few-shot graph-transformation tasks (21 transformations; the model sees a few input-output graph pairs and a test input graph of 5 to 15 nodes encoded as an adjacency or incidence list), answer compared with the value computed from the true output graph; the OpenAI reasoning models ran one system prompt at medium reasoning effort, the other models four system-prompt variants; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-589
2O4 Mini89
3O3 Mini79
4O1 Mini65
5Qwen 3 8B62
6GPT-4.1 Nano59
7Qwen 3 1.7B59
8Qwen 3 14B56
9Qwen 3 4B55
10Qwen 3 32B54
11OLMo 2 7B28
12DeepSeek R1 Distill Llama 8B27
13Llama 3.1 8B26
14Mistral-7B-v0.224
15Llama 3 8B16

Interactive version: theaggregate.ai/benchmark?slug=grapharc-output-graph-component-count · How It Works · Data refreshed daily, snapshot 2026-10-07.