CRUXEval: leaderboard

Code reasoning benchmark with 800 Python functions. Tests input prediction (given output, predict input) and output prediction (given input, predict output). Simple functions but tricky reasoning.

Metric: Output Prediction pass@1 (%). Source: crux-eval.github.io. Status: saturated. 43 models tracked.

Top models

#ModelScore
1GPT-4o70
2GPT-4 (0613)68.7
3GPT-4 Turbo67.7
4Claude 3 Opus65.8
5GPT-3.5 Turbo (0613)49.4
6starcoder2-15B47.1
7Mixtral 8x7B40.5
8starcoder2-7B36
9Mistral 7B34.3
10starcoder2-3B34.2
11Phi-233.5
12Phi-1.527.5
13Phi-121.7

Interactive version: theaggregate.ai/benchmark?slug=cruxeval · How It Works · Data refreshed daily, snapshot 2026-09-05.