PaperMind - Multimodal Grounding: leaderboard

Metric: Judge score for writing the caption of a figure from a real paper given the paper's introduction, the original caption as reference; rated by a GPT-4o judge on a 5-point scale (1-5) against the reference, averaged over the seven scientific domains; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro2.39
2GPT-4o Mini2.18
3Claude 3.5 Sonnet2.14
4Qwen 3 VL 4B Instruct2.11
5Gemma 3 4B (IT)1.89
6Claude 3 Haiku1.87

Interactive version: theaggregate.ai/benchmark?slug=papermind-multimodal-grounding · How It Works · Data refreshed daily, snapshot 2026-10-07.