DialBGM (One-Line Summary): leaderboard

Metric: Hit@1 (%, times 100): share of dialogues where the model's top-scored clip is the human rank-1 clip (chance 25), with a one-sentence GPT-4o summary of the dialogue as input, on DialBGM, 1,200 everyday multi-turn dialogues each paired with four candidate music clips ranked by human annotators for suitability as background music, the model scoring each audio clip; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro34.74
2Gemini 2.5 Flash32.28
3Phi-4 Multimodal Instruct26.22

Interactive version: theaggregate.ai/benchmark?slug=dialbgm-one-line-summary · How It Works · Data refreshed daily, snapshot 2026-10-07.