UrduMMLU (Urdu Prompt) - STEM: leaderboard

Metric: STEM (5,113 questions: biology, chemistry, computer science, general science, mathematics, physics) exact-match accuracy (%) on UrduMMLU's Urdu multiple-choice questions from Pakistani MCQ banks and public examinations, zero-shot, answer key parsed from the generation (unparsable, malformed or error outputs count as wrong); temperature 0 where available, 4,096 output tokens, Urdu instruction prompt. Source: arxiv.org. Saturation forecast: Estimated already saturated. 30 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash97.75
2GPT-5.497.4
3DeepSeek V4 Flash97.36
4Gemini 3.1 Flash Lite97.09
5Claude Sonnet 4.696.26
6Qwen 3.6 35B A3B96.15
7Gemma 4 31B (IT)93.86
8Llama 4 Maverick Instruct92.24
9Claude Haiku 4.591.92
10Qwen 3.6 27B91.12
11GPT-5.4 Mini88.25
12Gemma 4 26B A4B (IT)87.19
13Llama 4 Scout Instruct85.39
14Llama 3.3 70B Instruct78.39
15Qwen 3 8B74.05

Interactive version: theaggregate.ai/benchmark?slug=urdummlu-urdu-prompt-stem · How It Works · Data refreshed daily, snapshot 2026-09-29.