KoALa-Bench - KCSAT: leaderboard

Metric: Accuracy (%) on the long-form listening questions of the Korean College Scholastic Ability Test (2006-2012 recordings, multiple choice), Korean speech synthesized or recorded per item, clean audio (first value of each clean / noisy pair), mean over four prompts with greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash Lite81.18
2Qwen3 Omni 30B A3B Instruct80.59
3Voxtral-Mini-3B-250769.41
4Gemma 3n E4B (IT)33.53

Interactive version: theaggregate.ai/benchmark?slug=koala-bench-kcsat · How It Works · Data refreshed daily, snapshot 2026-10-07.