Phun-Bench - Contextual Homophone Recognition: leaderboard
Metric: Recognition rate (%): share of homophonic pun sentences (1,684) for which the model both finds the pun word and recovers the intended alternative word; thinking models run on 300 sampled items per setting, DeepSeek-V3 and GPT-4o on 500, the others on the full set; recommended sampling temperatures; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek R1 0528 | 81.3 |
| 2 | DeepSeek V3 (0324) | 78.2 |
| 3 | GPT-4o (2024-11-20) | 71.2 |
| 4 | Qwen 3 8B (Thinking) | 42.2 |
| 5 | Qwen 3 8B (Non-reasoning) | 34.5 |
| 6 | GLM-4 9B Chat | 25.6 |
| 7 | Llama 3.2 3B Instruct | 15.2 |
Interactive version: theaggregate.ai/benchmark?slug=phun-bench-contextual-homophone-recognition · How It Works · Data refreshed daily, snapshot 2026-09-29.