GPT Audio Mini: benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1609 ± 28, rank #427 of 1605 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| IHBench - Recovery Quality - Correction | 72 | Recovery pass rate (%; 25 user corrections, every type-speci | 82.7 |
| AOR-Bench | 20.2 | Over-refusal rate (%; refusals of 500 pseudo-harmful audio q | 81.8 |
| IHBench - Recovery Quality - Pushback | 71 | Recovery pass rate (%; 105 pushbacks, every type-specific cr | 69.2 |
| IHBench - Recovery Quality - Normal | 77 | Recovery pass rate (%; 81 normal cut-ins, every type-specifi | 59.6 |
| MSI-Bench - English Constraint Prioritization | 38.5 | All-pass rate (%; the 96 constraint-prioritization cases, wh | 57.1 |
| MSI-Bench - Mandarin Scope Tracking | 15.6 | All-pass rate (%; the 96 scope-tracking cases, where group-w | 57.1 |
| MSI-Bench - English Background Speech Retrieval | 7.3 | All-pass rate (%; the 96 background-speech-retrieval cases, | 53.6 |
| MSI-Bench - English Scope Tracking | 15.6 | All-pass rate (%; the 96 scope-tracking cases, where group-w | 53.6 |
| MSI-Bench - English Sequential Constraints | 20.8 | All-pass rate (%; the 96 sequential-constraint cases, where | 50 |
| MSI-Bench - Mandarin | 16.5 | All-pass rate (%; all 576 Mandarin test cases over six multi | 50 |
| MSI-Bench - Mandarin Constraint Prioritization | 27.1 | All-pass rate (%; the 96 constraint-prioritization cases, wh | 50 |
| MSI-Bench - Mandarin Sequential Constraints | 7.3 | All-pass rate (%; the 96 sequential-constraint cases, where | 50 |
Interactive version: theaggregate.ai/model?slug=gpt-audio-mini · How It Works · Data refreshed daily, snapshot 2026-09-26.