Pak3H - PakAlpaca (Urdu): leaderboard

Metric: Win rate (%; AlpacaEval pairwise preference of the model response over the language-matched reference response on all 4,445 human-contextualized Urdu PakAlpaca instructions, judged by GPT-4o-mini; temperature 0.7, 2,048 new tokens). Source: arxiv.org. Saturation forecast: Around August 2028. 3 models tracked.

Top models

#ModelScore
1Llama 3 70B Instruct36.2
2DeepSeek V3.229.26
3GPT-4o Mini13.19

Interactive version: theaggregate.ai/benchmark?slug=pak3h-pakalpaca-urdu · How It Works · Data refreshed daily, snapshot 2026-09-29.