MMLU-by-task - Elementary Mathematics — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 1257 models tracked.

Top models

#ModelScore
1falcon-180B48.94
2Llama 2 70B Chat45.5
3StableBeluga243.92
4Llama 2 70B Base43.39
5Llama 2 70B Chat (HF)41.01
6WizardCoder-Python-34B-V1.041.01
7LLaMA-65B40.48
8Llama 2 70B Chat GPTQ39.95
9internlm-20B Chat37.83
10vicuna-33B-v1.337.57
11Mistral-7B-v0.137.3
12LLaMA-30B37.3
13OpenHermes-13B35.98
14trurl-2-13B-academic35.45
15Llama 2 13B Chat Base34.13

Interactive version: theaggregate.ai/benchmark?slug=mmlu-by-task-elementary-mathematics · How the rankings work · Data refreshed daily, snapshot 2026-07-22.