CRUXEval — leaderboard

Code reasoning benchmark with 800 Python functions. Tests input prediction (given output, predict input) and output prediction (given input, predict output). Simple functions but tricky reasoning.

Metric: Output Prediction pass@1 (%). Source: crux-eval.github.io. Status: saturated. 43 models tracked.

Top models

#ModelScore
1GPT-477.1
2GPT-4o70
3GPT-4 (0613)68.7
4GPT-4 Turbo67.7
5Claude 3 Opus65.8
6GPT-3.5 Turbo59
7GPT-3.5 Turbo (0613)49.4
8starcoder2-15B47.1
9Mixtral 8x7B40.5
10starcoder2-7B36
11Mistral 7B34.3
12starcoder2-3B34.2
13Phi-233.5
14Phi-1.527.5
15Phi-121.7

Interactive version: theaggregate.ai/benchmark?slug=cruxeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.