YOMI-Bench - Rhyme Generation: leaderboard

Metric: Accuracy (%; generated hiragana word whose vowel sequence exactly matches the input word under a fixed vowel conversion table; 120 items; mean over five paraphrased prompts; accuracy as a fraction x100). Source: arxiv.org. Saturation forecast: Around May 2027. 10 models tracked.

Top models

#ModelScore
1Llama-3.1-Swallow-8B-Instruct-v0.584.8
2GPT-578
3Claude Sonnet 4.572.8
4Ministral-8B-Instruct-241065.19
5GPT-4o59.8
6Gemini 2.5 Flash59.2
7Mistral Medium 3.146.79

Interactive version: theaggregate.ai/benchmark?slug=yomi-bench-rhyme-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.