IndicKLAR (Urdu): leaderboard

Metric: Accuracy (%; relaxed prefix match of the answer to 2,619 factual queries in Urdu script, machine-translated, not speaker-verified; direct 3-shot prompting with English examples, greedy decoding). Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Gemma 3 12B (IT)68.8
2Qwen 2.5 14B Instruct59
3Llama 3.1 8B Instruct57
4Qwen 2.5 7B Instruct48.5
5Gemma 3 4B (IT)48.3
6Llama 3.2 3B Instruct40.7
7Llama 3.2 1B Instruct26.5
8Qwen 2.5 1.5B Instruct25.1
9Gemma 3 1B (IT)13.7

Interactive version: theaggregate.ai/benchmark?slug=indicklar-urdu · How It Works · Data refreshed daily, snapshot 2026-09-26.