RubricEval — leaderboard
Instruction-following evaluation framework for open-ended tasks using example-specific rubrics and scalable model grading.
Metric: Mean Score (1-5). Source: huggingface.co. Status: saturation imminent. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 | 3.18 |
| 2 | GPT-4 Turbo | 3.1 |
| 3 | Gemini 1.5 Pro | 3.06 |
| 4 | Gemini 1.5 Flash | 2.98 |
| 5 | Llama 3 70B | 2.9 |
| 6 | Claude 3 Opus | 2.86 |
| 7 | Claude 3 Sonnet | 2.79 |
| 8 | Claude 3 Haiku | 2.73 |
| 9 | Gemini 1.0 Pro | 2.56 |
| 10 | Llama 3 8B | 2.56 |
| 11 | GPT-3.5 Turbo | 2.52 |
| 12 | gemma-7B | 2.14 |
| 13 | gemma-2B | 1.83 |
Interactive version: theaggregate.ai/benchmark?slug=rubriceval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.