BuzzBench (Humour) — leaderboard

Humour analysis benchmark testing how well LLMs can understand and analyze comedy - a challenging component of emotional intelligence.

Metric: Humour Score (0-100). Source: eqbench.com. Status: saturation imminent. 26 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview 03-25)71.09
2Claude 3.7 Sonnet (20250219)68.2
3GPT-4.167.1
4DeepSeek R162.73
5Claude 3.5 Sonnet61.94
6GPT-4o ChatGPT59.48
7O159.28
8Gemini 1.5 Pro54.9
9Gemini 2.0 Flash (001)54.39
10Claude 3.5 Haiku (20241022)53.34
11Grok 2 (1212)51.76
12DeepSeek V351.63
13Mistral Large 2 (Nov) Instruct (2411)47.48
14O1 Mini46.18
15Llama 3.1 405B Instruct45.94

Interactive version: theaggregate.ai/benchmark?slug=buzzbench-humour · How the rankings work · Data refreshed daily, snapshot 2026-07-22.