VMLU - Humanities — leaderboard

Metric: Accuracy (%). Source: vmlu.ai. 25 models tracked.

Top models

#ModelScore
1QwQ-32B71.78
2Llama 3 70B68.74
3GPT-466.14
4GPT-4o Mini65.03
5QwQ 32B-Preview63.32
6Gemma 2 9B (IT)60.8
7Qwen 2.5 7B Instruct58.3
8gemma-7B (IT)43.39
9Phi-3-small-8k-instruct42.32
10Phi-3-small-128k-instruct41.78
11Qwen-7B34.15
12Qwen 2 7B Instruct33.13
13gemma-2B (IT)31.01
14falcon-7B26.72
15bloom-1b726.34

Interactive version: theaggregate.ai/benchmark?slug=vmlu-humanities · How the rankings work · Data refreshed daily, snapshot 2026-07-22.