ARC Challenge (AI2): leaderboard

AI2 Reasoning Challenge: 7,787 science exam questions at grade-school level. The 'Challenge' set contains questions that simple retrieval and co-occurrence methods fail on.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 77 models tracked.

Top models

#ModelScore
1DeepSeek V395.3
2Llama 3.1 405B95.3
3Qwen 2.5 72B94.5
4DeepSeek V292.2
5Phi-3-medium-128k-instruct91.6
6Phi-3-small-8k-instruct90.7
7GPT-3.5 Turbo (1106)87.4
8Mixtral 8x7B (v0.1)87.3
9Claude Instant 1.286.3
10Claude Instant 1.185.7
11text-davinci-00285.2
12Phi-3 Mini 4K Instruct84.9
13Qwen-14B84.4
14Llama 3 8B Instruct82.8
15Mistral-7B-v0.178.6

Interactive version: theaggregate.ai/benchmark?slug=arc-challenge-ai2 · How It Works · Data refreshed daily, snapshot 2026-09-05.