MMLU-by-task - Moral Disputes — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 1257 models tracked.

Top models

#ModelScore
1falcon-180B80.06
2Llama 2 70B Base77.75
3StableBeluga277.46
4Llama 2 70B Chat75.72
5LLaMA-65B73.7
6Llama 2 70B Chat (HF)71.68
7Llama 2 70B Chat GPTQ71.68
8Mistral-7B-v0.171.1
9vicuna-33B-v1.368.5
10LLaMA-30B67.05
11OpenHermes-13B64.74
12internlm-20B Chat64.74
13MythoMax-L2-13B62.43
14Llama 2 13B Chat Base61.27
15falcon-40B Instruct61.27

Interactive version: theaggregate.ai/benchmark?slug=mmlu-by-task-moral-disputes · How the rankings work · Data refreshed daily, snapshot 2026-07-22.