MEGA-Bench Task - Graph Interpretation — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro92.8
2Claude 3.5 Sonnet (20240620)88.3
3Claude 3.5 Sonnet (20241022)86.2
4Gemini 2.0 Flash (Preview)84.1
5GPT-4o Mini83.8
6GPT-4o83.1
7Gemini 1.5 Pro (002)82.4
8InternVL3-78B82.1
9Gemma 3 27B (IT)80.3
10Llama 4 Scout Base80.3
11Gemini 1.5 Flash (002)79
12InternVL2.5-78B78.6
13InternVL3-38B77.6
14Qwen 2 VL 72B76.2
15InternVL3-14B75.5

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-graph-interpretation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.