GPT Audio: benchmark results

Provider: OpenAI. Access: API.

Unified ELO 1754 ± 27, rank #146 of 1605 rated models, from 27 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
DualEvasion - Vocal Confidence58.5Macro F1 (%; zero-shot classification of the response audio 100
IHBench - Recovery Quality - Correction76Recovery pass rate (%; 25 user corrections, every type-speci96.2
IHBench - Recovery Quality - Pushback74Recovery pass rate (%; 105 pushbacks, every type-specific cr94.2
MSI-Bench - English Background Speech Retrieval14.6All-pass rate (%; the 96 background-speech-retrieval cases, 89.3
IHBench - Task Fulfillment64.4Win rate against GPT-4o Audio (%; share of 428 interruption 88
IHBench - Recovery Quality - Normal82Recovery pass rate (%; 81 normal cut-ins, every type-specifi86.5
MSI-Bench - Mandarin Selective Disclosure39.6All-pass rate (%; the 96 selective-disclosure cases, where e75
MSI-Bench - English Sequential Constraints44.8All-pass rate (%; the 96 sequential-constraint cases, where 71.4
MSI-Bench - Mandarin Constraint Prioritization49All-pass rate (%; the 96 constraint-prioritization cases, wh71.4
MSI-Bench - Mandarin Speaker Authority74All-pass rate (%; the 96 speaker-authority cases, where an u71.4
MSI-Bench - Mandarin Sequential Constraints44.8All-pass rate (%; the 96 sequential-constraint cases, where 67.9
MSI-Bench - English43.9All-pass rate (%; all 576 English test cases over six multi-64.3

Interactive version: theaggregate.ai/model?slug=gpt-audio · How It Works · Data refreshed daily, snapshot 2026-09-26.