LOL Bench - Joke Explanation: leaderboard
Blind-graded benchmark of how well language models explain jokes, scored as a mean grade over a fixed item pool with three samples per item.
Metric: Mean Grade (%). Source: www.lolbench.lol. Status: saturated. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 | 92.17 |
| 2 | Qwen 3.8 Max | 91.67 |
| 3 | Grok 4.6 | 90.6 |
| 4 | Qwen 3.8 Flash | 90.56 |
| 5 | GLM-5.3 Flash | 90.45 |
| 6 | GLM-5.3 | 90.38 |
| 7 | GPT-5.6 Pro Sol | 89.11 |
| 8 | Muse Spark 1.2 | 88.22 |
| 9 | Gemini 3.1 Pro (Preview) | 88.05 |
| 10 | DeepSeek V4 Pro | 87.26 |
| 11 | Hy3 | 86.74 |
| 12 | MiMo-V2.5-Pro | 83.73 |
| 13 | MiMo-V2.5 | 83.65 |
Interactive version: theaggregate.ai/benchmark?slug=lol-bench-joke-explanation · How It Works · Data refreshed daily, snapshot 2026-09-19.