MEGA-Bench — leaderboard

TIGER-Lab multimodal evaluation with 500+ tasks across vision-language understanding, generation, and reasoning with 44 models.

Metric: Overall Score (%). Source: huggingface.co. Status: saturation imminent. 51 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro64.7
2O158
3Gemini 2.0 Flash (Preview)55.3
4Claude 3.5 Sonnet (20241022)54.3
5GPT-4o54.2
6Claude 3.5 Sonnet (20240620)52.1
7Gemini 1.5 Pro (002)49.6
8InternVL3-78B47.1
9Qwen 2 VL 72B46.8
10InternVL3-38B45.8
11InternVL2.5-78B45.6
12Gemini 1.5 Flash (002)43.8
13Gemma 3 27B (IT)43.7
14GPT-4o Mini43.1
15InternVL3-14B42.5

Interactive version: theaggregate.ai/benchmark?slug=mega-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.