MEGA-Bench Task - Visual Prediction Rater Semantic Segmentation — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1GPT-4o60.4
2Gemini 1.5 Pro (002)54.2
3InternVL3-14B52.1
4Gemini 2.0 Flash (Preview)52.1
5InternVL3-38B50
6Gemini 2.5 Pro45.8
7Claude 3.5 Sonnet (20241022)45.8
8Gemma 3 12B (IT)43.8
9Gemma 3 27B (IT)41.7
10InternVL3-8B41.7
11Claude 3.5 Sonnet (20240620)41.7
12GPT-4o Mini39.6
13Gemini 1.5 Flash (002)39.6
14Gemma 3 4B (IT)31.2
15Qwen 2 VL 72B27.1

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-visual-prediction-rater-semantic-segmentation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.