AppTek Call-Center Dialogues (Manual Segmentation): leaderboard
Metric: Word error rate (%, averaged over the 14 accents) on AppTek Call-Center Dialogues (128.6 hours of spontaneous role-played English call-center conversations, 14 accents), default inference settings, scored per session after Open ASR Leaderboard text normalization, audio segmented with manual (human) segment boundaries; lower is better. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Phi-4 Multimodal Instruct | 9.2 |
| 2 | Whisper Large-v3 | 10.7 |
Interactive version: theaggregate.ai/benchmark?slug=apptek-call-center-dialogues-manual-segmentation · How It Works · Data refreshed daily, snapshot 2026-10-07.