MEGA-Bench - Output Format: Contextual Formatted Text: leaderboard

Metric: Average Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview 03-25)60.7
2GPT-4o53.9
3Claude 3.5 Sonnet (20241022)51.9
4Claude 3.5 Sonnet (20240620)50.7
5Gemini 1.5 Pro (002)44.9
6InternVL3-78B44.3
7Qwen 2 VL 72B43.6
8InternVL2.5-78B42.3
9Gemma 3 27B (IT)41.2
10InternVL3-38B41.2
11GPT-4o Mini41.2
12Gemini 1.5 Flash (002)38.7
13InternVL3-8B36.1
14Gemma 3 12B (IT)35.7
15Llama 4 Scout Base35.2

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-output-format-contextual-formatted-text · How It Works · Data refreshed daily, snapshot 2026-09-05.