LilyBench - Understanding: leaderboard

Metric: Macro average (0-1 scaled to %) over the eight Mutopia understanding tasks adapted from ABC-Eval to raw LilyPond text (bar count exact match, metadata QA, bar sequencing penalised Kendall tau, next-bar prediction, metadata prediction, music captioning, composer and genre recognition; 4-way multiple choice where applicable, chance 25); greedy decoding, at most 20 new tokens, no chain of thought; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1Phi-465.1
2Qwen 2.5 Coder 14B64.2

Interactive version: theaggregate.ai/benchmark?slug=lilybench-understanding · How It Works · Data refreshed daily, snapshot 2026-09-29.