MEGA-Bench Task - Code Translation Python — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1GPT-4o64.6
2Claude 3.5 Sonnet (20240620)64.6
3Claude 3.5 Sonnet (20241022)60.4
4Gemini 2.0 Flash (Preview)54.2
5Gemma 3 27B (IT)50
6InternVL3-38B50
7Gemini 2.5 Pro47.9
8InternVL3-14B45.8
9Gemini 1.5 Pro (002)45.8
10Gemma 3 12B (IT)43.8
11InternVL3-78B43.8
12Qwen 2 VL 72B41.7
13Gemini 1.5 Flash (002)41.7
14InternVL2.5-78B41.7
15InternVL3-8B35.4

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-code-translation-python · How the rankings work · Data refreshed daily, snapshot 2026-07-22.