MEGA-Bench Task - Visual Prediction Rater Panoptic Segmentation — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro88.1
2Claude 3.5 Sonnet (20241022)81
3Claude 3.5 Sonnet (20240620)59.5
4Gemini 2.0 Flash (Preview)54.8
5Gemini 1.5 Pro (002)50
6GPT-4o47.6
7InternVL3-78B42.9
8Gemma 3 27B (IT)38.1
9InternVL3-14B26.2
10Gemini 1.5 Flash (002)26.2
11Llama 4 Scout Base21.4
12InternVL3-38B19
13Gemma 3 12B (IT)14.3
14MiniCPM-V-2.614.3
15Gemma 3 4B (IT)9.5

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-visual-prediction-rater-panoptic-segmentation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.