GPT Audio Mini: benchmark results

Provider: OpenAI. Access: API.

Unified ELO 1609 ± 28, rank #427 of 1605 rated models, from 26 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
IHBench - Recovery Quality - Correction72Recovery pass rate (%; 25 user corrections, every type-speci82.7
AOR-Bench20.2Over-refusal rate (%; refusals of 500 pseudo-harmful audio q81.8
IHBench - Recovery Quality - Pushback71Recovery pass rate (%; 105 pushbacks, every type-specific cr69.2
IHBench - Recovery Quality - Normal77Recovery pass rate (%; 81 normal cut-ins, every type-specifi59.6
MSI-Bench - English Constraint Prioritization38.5All-pass rate (%; the 96 constraint-prioritization cases, wh57.1
MSI-Bench - Mandarin Scope Tracking15.6All-pass rate (%; the 96 scope-tracking cases, where group-w57.1
MSI-Bench - English Background Speech Retrieval7.3All-pass rate (%; the 96 background-speech-retrieval cases, 53.6
MSI-Bench - English Scope Tracking15.6All-pass rate (%; the 96 scope-tracking cases, where group-w53.6
MSI-Bench - English Sequential Constraints20.8All-pass rate (%; the 96 sequential-constraint cases, where 50
MSI-Bench - Mandarin16.5All-pass rate (%; all 576 Mandarin test cases over six multi50
MSI-Bench - Mandarin Constraint Prioritization27.1All-pass rate (%; the 96 constraint-prioritization cases, wh50
MSI-Bench - Mandarin Sequential Constraints7.3All-pass rate (%; the 96 sequential-constraint cases, where 50

Interactive version: theaggregate.ai/model?slug=gpt-audio-mini · How It Works · Data refreshed daily, snapshot 2026-09-26.