GPT Realtime 2: benchmark results

Provider: OpenAI. Access: API.

Unified ELO 1761 ± 22, rank #115 of 1608 rated models, from 24 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
IHBench - Task Fulfillment72.8Win rate against GPT-4o Audio (%; share of 428 interruption 100
RedVox1.1Unsafe response rate (%; share of responses a GPT-5.5 judge 100
RedVox - English0Unsafe response rate (whole percent, truncated as printed; s100
RedVox - French1Unsafe response rate (whole percent, truncated as printed; s100
RedVox - German2Unsafe response rate (whole percent, truncated as printed; s100
RedVox - Italian1Unsafe response rate (whole percent, truncated as printed; s100
RedVox - Spanish1Unsafe response rate (whole percent, truncated as printed; s92.9
IHBench - Recovery Quality - Correction73Recovery pass rate (%; 25 user corrections, every type-speci88.5
AudioMC48.45Score (self-reported)85.2
IHBench - Recovery Quality - Pushback72Recovery pass rate (%; 105 pushbacks, every type-specific cr82.7
VAmoS Bench67Task Completion (%)77.8
ODU-Bench (Audio-Only) - Key-Point Coverage63.4Key-point hit rate (%; share of reference key points on stru69.2

Interactive version: theaggregate.ai/model?slug=gpt-realtime-2 · How It Works · Data refreshed daily, snapshot 2026-10-06.