HumorRank - SemEval-2026 MWAHAHA: leaderboard
Metric: Bradley-Terry rating (open scale, mean about 1000) from a full round-robin of pairwise joke comparisons on SemEval-2026 MWAHAHA, 10,800 judgments, judged by Llama 3.3 70B Instruct; higher is better. Source: arxiv.org. Saturation forecast: Around October 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 1307.5 |
| 2 | Kimi K2 | 1156.9 |
| 3 | Gemini 2.5 Pro | 1115.1 |
| 4 | Claude 3.5 Haiku | 1037.5 |
| 5 | GPT-OSS-120B | 1015 |
| 6 | Qwen 3 32B | 976.9 |
| 7 | Llama 3.3 70B | 761 |
| 8 | Qwen 2.5 7B Instruct | 537.4 |
Interactive version: theaggregate.ai/benchmark?slug=humorrank-semeval-2026-mwahaha · How It Works · Data refreshed daily, snapshot 2026-10-07.