ENEM Item Difficulty (Point-Based Prompt): leaderboard
Metric: Spearman rank correlation (-1 to 1) between the model's 1-10 difficulty rating and the official IRT-derived difficulty of 1,031 released ENEM items (2017-2022, four subject areas), zero-shot with the point-based rubric prompt (the paper's pre-declared primary contrast); items with no parsable rating are left out; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Flash (Preview 04-17) | 0.37 | #230 |
| 2 | Gemini 2.0 Flash | 0.32 | #331 |
| 3 | GPT-4o (2024-08-06) | 0.32 | #326 |
| 4 | O3 (2025-04-16) | 0.31 | #117 |
| 5 | Mistral Large 2 (Nov) Instruct (2411) | 0.23 | #458 |
| 6 | DeepSeek R1 | 0.2 | #245 |
| 7 | Qwen 3 14B | 0.12 | #524 |
| 8 | Phi-4 | 0.09 | #701 |
| 9 | Llama 3.2 3B Instruct | -0.03 | #1321 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=enem-item-difficulty-point-based-prompt · How It Works · Data refreshed daily, snapshot 2026-10-11.