UrduMMLU (Urdu Prompt) - General Knowledge: leaderboard

Metric: General knowledge (1,362 questions, the Other domain) exact-match accuracy (%) on UrduMMLU's Urdu multiple-choice questions from Pakistani MCQ banks and public examinations, zero-shot, answer key parsed from the generation (unparsable, malformed or error outputs count as wrong); temperature 0 where available, 4,096 output tokens, Urdu instruction prompt. Source: arxiv.org. Saturation forecast: Around December 2026. 30 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash91.85
2Claude Sonnet 4.686.05
3Gemini 3.1 Flash Lite85.76
4DeepSeek V4 Flash84.88
5GPT-5.483.92
6Llama 4 Maverick Instruct80.98
7Gemma 4 31B (IT)78.56
8Qwen 3.6 35B A3B77.53
9GPT-5.4 Mini75.26
10Claude Haiku 4.574.45
11Gemma 4 26B A4B (IT)71.95
12Llama 3.3 70B Instruct71.81
13Qwen 3.6 27B70.7
14Llama 4 Scout Instruct69.16
15Ministral-3-8B-Instruct-251257.12

Interactive version: theaggregate.ai/benchmark?slug=urdummlu-urdu-prompt-general-knowledge · How It Works · Data refreshed daily, snapshot 2026-09-29.