MEGA-Bench Task - Image Captioning With Additional Requirements — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20241022)94.3
2Gemini 2.5 Pro93.6
3Claude 3.5 Sonnet (20240620)93.6
4InternVL3-38B92.9
5InternVL3-14B92.9
6Gemini 2.0 Flash (Preview)92.9
7GPT-4o92.1
8InternVL3-78B92.1
9InternVL2.5-78B91.4
10Ovis2-8B90.7
11GPT-4o Mini90
12Gemma 3 12B (IT)89.3
13Gemma 3 27B (IT)88.6
14Gemini 1.5 Pro (002)88.6
15InternVL3-8B86.4

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-image-captioning-with-additional-requirements · How the rankings work · Data refreshed daily, snapshot 2026-07-22.