DialBGM (One-Line Summary) - Kendall Tau: leaderboard
Metric: Kendall tau-b (from -1 to 1) between the model's ranking of the four clips and the human ranking, averaged over dialogues, with a one-sentence GPT-4o summary of the dialogue as input, on DialBGM, 1,200 everyday multi-turn dialogues each paired with four candidate music clips ranked by human annotators for suitability as background music, the model scoring each audio clip; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 0.18 |
| 2 | Gemini 2.5 Flash | 0.18 |
| 3 | Phi-4 Multimodal Instruct | 0.02 |
Interactive version: theaggregate.ai/benchmark?slug=dialbgm-one-line-summary-kendall-tau · How It Works · Data refreshed daily, snapshot 2026-10-07.