ARC Challenge (AI2) — leaderboard

AI2 Reasoning Challenge: 7,787 science exam questions at grade-school level. The 'Challenge' set contains questions that simple retrieval and co-occurrence methods fail on.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 77 models tracked.

Top models

#ModelScore
1Human Expert100
2DeepSeek V395.3
3Llama 3.1 405B95.3
4Qwen 2.5 72B94.5
5DeepSeek V292.2
6Phi-3-medium-128k-instruct91.6
7Phi-3-small-8k-instruct90.7
8GPT-3.5 Turbo (1106)87.4
9Mixtral 8x7B (v0.1)87.3
10Claude Instant 1.286.3
11StableBeluga286.1
12Claude Instant 1.185.7
13text-davinci-00285.2
14Phi-3 Mini 4K Instruct84.9
15Qwen-14B84.4

Interactive version: theaggregate.ai/benchmark?slug=arc-challenge-ai2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.