GPT Realtime 2: benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1761 ± 22, rank #115 of 1608 rated models, from 24 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| IHBench - Task Fulfillment | 72.8 | Win rate against GPT-4o Audio (%; share of 428 interruption | 100 |
| RedVox | 1.1 | Unsafe response rate (%; share of responses a GPT-5.5 judge | 100 |
| RedVox - English | 0 | Unsafe response rate (whole percent, truncated as printed; s | 100 |
| RedVox - French | 1 | Unsafe response rate (whole percent, truncated as printed; s | 100 |
| RedVox - German | 2 | Unsafe response rate (whole percent, truncated as printed; s | 100 |
| RedVox - Italian | 1 | Unsafe response rate (whole percent, truncated as printed; s | 100 |
| RedVox - Spanish | 1 | Unsafe response rate (whole percent, truncated as printed; s | 92.9 |
| IHBench - Recovery Quality - Correction | 73 | Recovery pass rate (%; 25 user corrections, every type-speci | 88.5 |
| AudioMC | 48.45 | Score (self-reported) | 85.2 |
| IHBench - Recovery Quality - Pushback | 72 | Recovery pass rate (%; 105 pushbacks, every type-specific cr | 82.7 |
| VAmoS Bench | 67 | Task Completion (%) | 77.8 |
| ODU-Bench (Audio-Only) - Key-Point Coverage | 63.4 | Key-point hit rate (%; share of reference key points on stru | 69.2 |
Interactive version: theaggregate.ai/model?slug=gpt-realtime-2 · How It Works · Data refreshed daily, snapshot 2026-10-06.