MUNIChus - Chinese: leaderboard

Metric: CIDEr of the zero-shot Chinese news image caption generated from the image and its news article in the target language, scored by CIDEr against the journalist's caption on the MUNIChus test split (Chinese and Japanese segmented with Jieba and MeCab), printed times 100; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 2.5 VL 7B Instruct28.68#643
2GPT-4o22.29#333
3aya-vision-8B17.45#1094
4Llama 3.2 11B Instruct8.69#1112

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=munichus-chinese · How It Works · Data refreshed daily, snapshot 2026-10-11.