EchoChain: leaderboard

Metric: Mean pass rate (%): share of conversations in which the post-interruption response satisfies every rubric criterion, over 200 human-screened interrupted conversations in which the user barges in mid-response with new task information (pre-generated cloned-voice audio, fixed barge-in timing), each model response graded blind by human annotators against an instance-specific rubric; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1Grok Voice Agent48.5
2GPT-realtime-2025-08-2845
3Amazon Nova 2 Sonic26.5
4Gemini Live 2.5 Flash Native Audio16.5

Interactive version: theaggregate.ai/benchmark?slug=echochain · How It Works · Data refreshed daily, snapshot 2026-10-07.