Phun-Bench - Homophone Recall: leaderboard
Metric: Accuracy (%) of recovering a four-character chengyu from an automatically generated homophone of it (1,098 items); thinking models run on 300 sampled items per setting, DeepSeek-V3 and GPT-4o on 500, the others on the full set; recommended sampling temperatures; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek R1 0528 | 97.7 |
| 2 | DeepSeek V3 (0324) | 66.4 |
| 3 | GPT-4o (2024-11-20) | 61.8 |
| 4 | Qwen 3 8B (Thinking) | 56.3 |
| 5 | Qwen 3 8B (Non-reasoning) | 15.2 |
| 6 | GLM-4 9B Chat | 12.4 |
| 7 | Llama 3.2 3B Instruct | 0.7 |
Interactive version: theaggregate.ai/benchmark?slug=phun-bench-homophone-recall · How It Works · Data refreshed daily, snapshot 2026-09-29.