LOL Bench - Joke Explanation: leaderboard

Blind-graded benchmark of how well language models explain jokes, scored as a mean grade over a fixed item pool with three samples per item.

Metric: Mean Grade (%). Source: www.lolbench.lol. Status: saturated. 13 models tracked.

Top models

#ModelScore
1Claude Opus 592.17
2Qwen 3.8 Max91.67
3Grok 4.690.6
4Qwen 3.8 Flash90.56
5GLM-5.3 Flash90.45
6GLM-5.390.38
7GPT-5.6 Pro Sol89.11
8Muse Spark 1.288.22
9Gemini 3.1 Pro (Preview)88.05
10DeepSeek V4 Pro87.26
11Hy386.74
12MiMo-V2.5-Pro83.73
13MiMo-V2.583.65

Interactive version: theaggregate.ai/benchmark?slug=lol-bench-joke-explanation · How It Works · Data refreshed daily, snapshot 2026-09-19.