tau-Voice (Realistic) - Retail: leaderboard
Metric: pass@1 (%): share of tasks whose final database state matches the goal (agent communications checked by an LLM evaluator), full-duplex voice agent talking to a GPT-4.1 voice user simulator (ElevenLabs v3 speech) under the Realistic condition (diverse accents, background and burst noise, frame drops, muffling, LLM-driven interruptions, backchannels and non-directed speech), on the 114 retail tasks (returns, exchanges, cancellations, order changes); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT Realtime 1.5 | 45 | |
| 2 | Grok Voice Agent | 39 | |
| 3 | Gemini Live 2.5 Flash Native Audio | 30 |
Interactive version: theaggregate.ai/benchmark?slug=tau-voice-realistic-retail · How It Works · Data refreshed daily, snapshot 2026-10-11.