AudioProcessBench - Semantic Errors: leaderboard

Metric: Type-conditioned PRMScore (0-100): detection of steps with semantic errors (misreading speech content or meaning) while keeping correct steps, other error types left out; 3,872 reasoning traces (23,497 steps, 41% erroneous) from six audio and omni models on MMAR, MMSU and MMAU-Pro questions, steps labelled by two LLM annotators with human review; few-shot critic prompt with three annotated example chains, temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 11 models tracked.

Top models

#ModelScore
1Gemini 3 Flash69
2Gemma 3n E4B56.6
3Gemma 4 E4B55.2
4Phi-4 Multimodal Instruct50.3
5Gemma 3n E2B47.6
6Qwen2.5-Omni-7B44
7Gemma 4 E2B38.9

Interactive version: theaggregate.ai/benchmark?slug=audioprocessbench-semantic-errors · How It Works · Data refreshed daily, snapshot 2026-09-29.