HumorRank - SemEval-2026 MWAHAHA: leaderboard

Metric: Bradley-Terry rating (open scale, mean about 1000) from a full round-robin of pairwise joke comparisons on SemEval-2026 MWAHAHA, 10,800 judgments, judged by Llama 3.3 70B Instruct; higher is better. Source: arxiv.org. Saturation forecast: Around October 2026. 9 models tracked.

Top models

#ModelScore
1GPT-51307.5
2Kimi K21156.9
3Gemini 2.5 Pro1115.1
4Claude 3.5 Haiku1037.5
5GPT-OSS-120B1015
6Qwen 3 32B976.9
7Llama 3.3 70B761
8Qwen 2.5 7B Instruct537.4

Interactive version: theaggregate.ai/benchmark?slug=humorrank-semeval-2026-mwahaha · How It Works · Data refreshed daily, snapshot 2026-10-07.