Jamendo-MT-QA - Yes-No Questions: leaderboard
Metric: Accuracy (%) on comparative yes/no questions about two music tracks, rule-based answer matching; the 2,010-pair random subset on which every multi-audio model ran; each model hears both tracks and the question end to end; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o-mini-Audio | 77.3 |
| 2 | GPT-4o-Audio | 69.6 |
| 3 | Qwen3-Omni | 59.7 |
| 4 | Jamendo-MT-QA Qwen2-Audio (checkpoint unspecified) | 34.4 |
Interactive version: theaggregate.ai/benchmark?slug=jamendo-mt-qa-yes-no-questions · How It Works · Data refreshed daily, snapshot 2026-10-07.