MUNIChus - English: leaderboard
Metric: CIDEr of the zero-shot English news image caption generated from the image and its news article in the target language, scored by CIDEr against the journalist's caption on the MUNIChus test split (Chinese and Japanese segmented with Jieba and MeCab), printed times 100; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-4o | 39.64 | #333 |
| 2 | Qwen 2.5 VL 7B Instruct | 32.69 | #643 |
| 3 | aya-vision-8B | 28.13 | #1094 |
| 4 | Llama 3.2 11B Instruct | 19.37 | #1112 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=munichus-english · How It Works · Data refreshed daily, snapshot 2026-10-11.