Phun-Bench - Contextual Homophone Recognition: leaderboard

Metric: Recognition rate (%): share of homophonic pun sentences (1,684) for which the model both finds the pun word and recovers the intended alternative word; thinking models run on 300 sampled items per setting, DeepSeek-V3 and GPT-4o on 500, the others on the full set; recommended sampling temperatures; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1DeepSeek R1 052881.3
2DeepSeek V3 (0324)78.2
3GPT-4o (2024-11-20)71.2
4Qwen 3 8B (Thinking)42.2
5Qwen 3 8B (Non-reasoning)34.5
6GLM-4 9B Chat25.6
7Llama 3.2 3B Instruct15.2

Interactive version: theaggregate.ai/benchmark?slug=phun-bench-contextual-homophone-recognition · How It Works · Data refreshed daily, snapshot 2026-09-29.