VocalCoachBench - Triplet Ranking: leaderboard
Metric: Pairwise accuracy (%; for three performances of the same song the model compares each pair, A vs B, A vs C and B vs C, and each preference is scored against the ranking induced by professional vocal trainers; random choice scores 50; main structured-prompt protocol, deterministic decoding where supported). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 Omni Plus | 74.5 |
| 2 | Gemini 3 Flash (Preview) | 71.2 |
Interactive version: theaggregate.ai/benchmark?slug=vocalcoachbench-triplet-ranking · How It Works · Data refreshed daily, snapshot 2026-09-26.