InstructEval - Problem Solving: leaderboard
Metric: Average (%, MMLU/BBH/DROP/CRASS/HumanEval). Source: declare-lab.github.io. 54 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | flan-ul2 | 51.6 |
| 2 | flan-t5-xxl | 50.8 |
| 3 | flan-t5-xl | 47.4 |
| 4 | flan-t5-large | 41 |
| 5 | alpaca-7B | 32.4 |
| 6 | flan-t5-base | 31.6 |
| 7 | starcoder | 27.5 |
| 8 | falcon-7B | 23.7 |
| 9 | falcon-7B Instruct | 22.2 |
Interactive version: theaggregate.ai/benchmark?slug=instructeval-problem-solving · How It Works · Data refreshed daily, snapshot 2026-09-05.