YOMI-Bench - Kanji Reading QA (Single Reading): leaderboard

Metric: Accuracy (%; yes/no judgement of whether a given reading is correct for a kanji in a word, 4-shot with two positive and two negative examples; 240 items on kanji with a single reading; mean over five paraphrased prompts; accuracy as a fraction x100). Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1GPT-4o100
2Gemini 2.5 Flash99.8
3GPT-599.7
4Mistral Medium 3.199.6
5Claude Sonnet 4.591.4
6Llama-3.1-Swallow-8B-Instruct-v0.576.2
7Ministral-8B-Instruct-241066.4

Interactive version: theaggregate.ai/benchmark?slug=yomi-bench-kanji-reading-qa-single-reading · How It Works · Data refreshed daily, snapshot 2026-09-29.