YOMI-Bench - Kanji Reading QA (Multiple Readings): leaderboard

Metric: Accuracy (%; yes/no judgement of whether a given reading is correct for a kanji in a word, 4-shot with two positive and two negative examples; 450 items on kanji with several readings; mean over five paraphrased prompts; accuracy as a fraction x100). Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1GPT-598.48
2Gemini 2.5 Flash98.17
3GPT-4o94.04
4Mistral Medium 3.187.37
5Llama-3.1-Swallow-8B-Instruct-v0.577.82
6Claude Sonnet 4.577.33
7Ministral-8B-Instruct-241063.55

Interactive version: theaggregate.ai/benchmark?slug=yomi-bench-kanji-reading-qa-multiple-readings · How It Works · Data refreshed daily, snapshot 2026-09-29.