Mobile-MMLU — leaderboard
Metric: Overall Accuracy (%). Source: huggingface.co. 22 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 2 9B (IT) | 75 |
| 2 | Qwen 2.5 7B Instruct | 74.9 |
| 3 | Yi-1.5-9B Chat | 72.7 |
| 4 | Ministral-8B-Instruct-2410 | 71.5 |
| 5 | Qwen 2.5 3B Instruct | 68.1 |
| 6 | Llama 3.1 8B Instruct | 66.9 |
| 7 | internlm2.5-7B-chat | 64.3 |
| 8 | Phi-3.5-mini-instruct | 63.7 |
| 9 | granite-3.1-8B-instruct | 60.8 |
| 10 | Yi-1.5-6B-Chat | 60.5 |
| 11 | EXAONE-3.5-2.4B-Instruct | 53.7 |
| 12 | Llama 3.2 3B Instruct | 50.2 |
| 13 | Qwen 2.5 1.5B Instruct | 49.7 |
| 14 | OLMo-2-1124-7B-Instruct | 49.6 |
| 15 | Gemma 2 2B (IT) | 38.9 |
Interactive version: theaggregate.ai/benchmark?slug=mobile-mmlu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.