OmniACBench - Pronunciation PER: leaderboard

Metric: Phoneme error rate as printed (human reference 1.21) of the generated speech, phonemes extracted by POWSM, against the reference phoneme sequence on the pronunciation instances, on OmniACBench's 3,559 synthetic instances (spoken instruction, text script and image; the model must read the script aloud with the delivery the context implies); lower is better. Source: arxiv.org. Saturation forecast: Around April 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1Qwen3 Omni 30B A3B Instruct7.4#362
2Qwen2.5-Omni-7B10.27#601

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=omniacbench-pronunciation-per · How It Works · Data refreshed daily, snapshot 2026-10-11.