MEGA-Bench - Output Format: Contextual Formatted Text — leaderboard

Metric: Average Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro60.7
2GPT-4o53.9
3Gemini 2.0 Flash (Preview)53
4Claude 3.5 Sonnet (20241022)51.9
5Claude 3.5 Sonnet (20240620)50.7
6Gemini 1.5 Pro (002)44.9
7InternVL3-78B44.3
8Qwen 2 VL 72B43.6
9InternVL2.5-78B42.3
10Gemma 3 27B (IT)41.2
11InternVL3-38B41.2
12GPT-4o Mini41.2
13InternVL3-14B39.5
14Gemini 1.5 Flash (002)38.7
15InternVL3-8B36.1

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-output-format-contextual-formatted-text · How the rankings work · Data refreshed daily, snapshot 2026-07-22.