ENEM Item Difficulty (Point-Based Prompt): leaderboard

Metric: Spearman rank correlation (-1 to 1) between the model's 1-10 difficulty rating and the official IRT-derived difficulty of 1,031 released ENEM items (2017-2022, four subject areas), zero-shot with the point-based rubric prompt (the paper's pre-declared primary contrast); items with no parsable rating are left out; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash (Preview 04-17)0.37#230
2Gemini 2.0 Flash0.32#331
3GPT-4o (2024-08-06)0.32#326
4O3 (2025-04-16)0.31#117
5Mistral Large 2 (Nov) Instruct (2411)0.23#458
6DeepSeek R10.2#245
7Qwen 3 14B0.12#524
8Phi-40.09#701
9Llama 3.2 3B Instruct-0.03#1321

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=enem-item-difficulty-point-based-prompt · How It Works · Data refreshed daily, snapshot 2026-10-11.