MUNIChus - French: leaderboard

Metric: CIDEr of the zero-shot French news image caption generated from the image and its news article in the target language, scored by CIDEr against the journalist's caption on the MUNIChus test split (Chinese and Japanese segmented with Jieba and MeCab), printed times 100; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScoreOverall rank
1aya-vision-8B14.71#1094
2Qwen 2.5 VL 7B Instruct14.08#643
3Llama 3.2 11B Instruct12.41#1112
4GPT-4o11.31#333

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=munichus-french · How It Works · Data refreshed daily, snapshot 2026-10-11.