MEGA-Bench Task - Counterfactual Arithmetic — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro100
2Gemma 3 27B (IT)100
3GPT-4o Mini92.9
4GPT-4o85.7
5InternVL3-8B85.7
6Gemini 1.5 Pro (002)85.7
7Llama 4 Scout Base85.7
8Gemini 2.0 Flash (Preview)85.7
9Claude 3.5 Sonnet (20240620)78.6
10Gemma 3 4B (IT)42.9
11InternVL3-14B42.9
12Qwen 2 VL 72B35.7
13Gemini 1.5 Flash (002)35.7
14Gemma 3 12B (IT)21.4
15Claude 3.5 Sonnet (20241022)14.3

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-counterfactual-arithmetic · How the rankings work · Data refreshed daily, snapshot 2026-07-22.