DialBGM (One-Line Summary) - Kendall Tau: leaderboard

Metric: Kendall tau-b (from -1 to 1) between the model's ranking of the four clips and the human ranking, averaged over dialogues, with a one-sentence GPT-4o summary of the dialogue as input, on DialBGM, 1,200 everyday multi-turn dialogues each paired with four candidate music clips ranked by human annotators for suitability as background music, the model scoring each audio clip; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro0.18
2Gemini 2.5 Flash0.18
3Phi-4 Multimodal Instruct0.02

Interactive version: theaggregate.ai/benchmark?slug=dialbgm-one-line-summary-kendall-tau · How It Works · Data refreshed daily, snapshot 2026-10-07.