MUNIChus - Urdu: leaderboard

Metric: CIDEr of the zero-shot Urdu news image caption generated from the image and its news article in the target language, scored by CIDEr against the journalist's caption on the MUNIChus test split (Chinese and Japanese segmented with Jieba and MeCab), printed times 100; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o26.19#333
2Qwen 2.5 VL 7B Instruct22.46#643
3Llama 3.2 11B Instruct15.61#1112
4aya-vision-8B12.84#1094

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=munichus-urdu · How It Works · Data refreshed daily, snapshot 2026-10-11.