UrduMMLU (Urdu Prompt) - Humanities: leaderboard

Metric: Humanities (10,982 questions: ethics, fine arts, Islamic studies, Urdu grammar, language and literature) exact-match accuracy (%) on UrduMMLU's Urdu multiple-choice questions from Pakistani MCQ banks and public examinations, zero-shot, answer key parsed from the generation (unparsable, malformed or error outputs count as wrong); temperature 0 where available, 4,096 output tokens, Urdu instruction prompt. Source: arxiv.org. Saturation forecast: Around December 2026. 30 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash85.49
2Gemini 3.1 Flash Lite74.57
3GPT-5.474.29
4Claude Sonnet 4.672.87
5DeepSeek V4 Flash67.3
6Llama 4 Maverick Instruct63.41
7Gemma 4 31B (IT)63.38
8GPT-5.4 Mini62.51
9Claude Haiku 4.559.42
10Qwen 3.6 35B A3B57.98
11Gemma 4 26B A4B (IT)57.87
12Llama 4 Scout Instruct56.69
13Llama 3.3 70B Instruct56.24
14Qwen 3.6 27B55.86
15Gemma 2 9B (IT)48.21

Interactive version: theaggregate.ai/benchmark?slug=urdummlu-urdu-prompt-humanities · How It Works · Data refreshed daily, snapshot 2026-09-29.