GPT Audio: benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1754 ± 27, rank #146 of 1605 rated models, from 27 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DualEvasion - Vocal Confidence | 58.5 | Macro F1 (%; zero-shot classification of the response audio | 100 |
| IHBench - Recovery Quality - Correction | 76 | Recovery pass rate (%; 25 user corrections, every type-speci | 96.2 |
| IHBench - Recovery Quality - Pushback | 74 | Recovery pass rate (%; 105 pushbacks, every type-specific cr | 94.2 |
| MSI-Bench - English Background Speech Retrieval | 14.6 | All-pass rate (%; the 96 background-speech-retrieval cases, | 89.3 |
| IHBench - Task Fulfillment | 64.4 | Win rate against GPT-4o Audio (%; share of 428 interruption | 88 |
| IHBench - Recovery Quality - Normal | 82 | Recovery pass rate (%; 81 normal cut-ins, every type-specifi | 86.5 |
| MSI-Bench - Mandarin Selective Disclosure | 39.6 | All-pass rate (%; the 96 selective-disclosure cases, where e | 75 |
| MSI-Bench - English Sequential Constraints | 44.8 | All-pass rate (%; the 96 sequential-constraint cases, where | 71.4 |
| MSI-Bench - Mandarin Constraint Prioritization | 49 | All-pass rate (%; the 96 constraint-prioritization cases, wh | 71.4 |
| MSI-Bench - Mandarin Speaker Authority | 74 | All-pass rate (%; the 96 speaker-authority cases, where an u | 71.4 |
| MSI-Bench - Mandarin Sequential Constraints | 44.8 | All-pass rate (%; the 96 sequential-constraint cases, where | 67.9 |
| MSI-Bench - English | 43.9 | All-pass rate (%; all 576 English test cases over six multi- | 64.3 |
Interactive version: theaggregate.ai/model?slug=gpt-audio · How It Works · Data refreshed daily, snapshot 2026-09-26.