MEGA-Bench Task - Code Execution — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro93.8
2Gemini 2.0 Flash (Preview)81.2
3GPT-4o75
4InternVL3-78B50
5GPT-4o Mini50
6Claude 3.5 Sonnet (20240620)50
7Gemma 3 27B (IT)43.8
8Gemini 1.5 Pro (002)43.8
9Gemini 1.5 Flash (002)43.8
10InternVL2.5-78B43.8
11InternVL3-14B37.5
12Claude 3.5 Sonnet (20241022)37.5
13Gemma 3 12B (IT)25
14InternVL2-8B25
15Qwen 2 VL 72B25

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-code-execution · How the rankings work · Data refreshed daily, snapshot 2026-07-22.