LilyBench - Understanding: leaderboard
Metric: Macro average (0-1 scaled to %) over the eight Mutopia understanding tasks adapted from ABC-Eval to raw LilyPond text (bar count exact match, metadata QA, bar sequencing penalised Kendall tau, next-bar prediction, metadata prediction, music captioning, composer and genre recognition; 4-way multiple choice where applicable, chance 25); greedy decoding, at most 20 new tokens, no chain of thought; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Phi-4 | 65.1 |
| 2 | Qwen 2.5 Coder 14B | 64.2 |
Interactive version: theaggregate.ai/benchmark?slug=lilybench-understanding · How It Works · Data refreshed daily, snapshot 2026-09-29.