tau-Voice (Clean) - Telecom: leaderboard

Metric: pass@1 (%): share of tasks whose final database state matches the goal (agent communications checked by an LLM evaluator), full-duplex voice agent talking to a GPT-4.1 voice user simulator (ElevenLabs v3 speech) under the Clean condition (clear American-accented speech, G.711 telephony, no noise or interruptions), on the 114 telecom tasks (plan changes, billing, activations, account changes); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScoreOverall rank
1Grok Voice Agent58
2GPT Realtime 1.528
3Gemini Live 2.5 Flash Native Audio20

Interactive version: theaggregate.ai/benchmark?slug=tau-voice-clean-telecom · How It Works · Data refreshed daily, snapshot 2026-10-11.