BuzzBench (Humour) — leaderboard
Humour analysis benchmark testing how well LLMs can understand and analyze comedy - a challenging component of emotional intelligence.
Metric: Humour Score (0-100). Source: eqbench.com. Status: saturation imminent. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro (Preview 03-25) | 71.09 |
| 2 | Claude 3.7 Sonnet (20250219) | 68.2 |
| 3 | GPT-4.1 | 67.1 |
| 4 | DeepSeek R1 | 62.73 |
| 5 | Claude 3.5 Sonnet | 61.94 |
| 6 | GPT-4o ChatGPT | 59.48 |
| 7 | O1 | 59.28 |
| 8 | Gemini 1.5 Pro | 54.9 |
| 9 | Gemini 2.0 Flash (001) | 54.39 |
| 10 | Claude 3.5 Haiku (20241022) | 53.34 |
| 11 | Grok 2 (1212) | 51.76 |
| 12 | DeepSeek V3 | 51.63 |
| 13 | Mistral Large 2 (Nov) Instruct (2411) | 47.48 |
| 14 | O1 Mini | 46.18 |
| 15 | Llama 3.1 405B Instruct | 45.94 |
Interactive version: theaggregate.ai/benchmark?slug=buzzbench-humour · How the rankings work · Data refreshed daily, snapshot 2026-07-22.