MEGA-Bench Task - Humor Explanation — leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro96
2GPT-4o86.7
3Gemini 1.5 Flash (002)85.3
4Gemini 2.0 Flash (Preview)85.3
5GPT-4o Mini82
6Gemini 1.5 Pro (002)80
7Gemma 3 27B (IT)79.3
8Claude 3.5 Sonnet (20241022)72.7
9Ovis2-8B72
10Gemma 3 12B (IT)70
11InternVL2.5-78B68.7
12Gemma 3 4B (IT)66.7
13InternVL3-14B59.3
14Claude 3.5 Sonnet (20240620)58.7
15InternVL3-8B55.3

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-humor-explanation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.