Jamendo-MT-QA - Short-Answer Questions: leaderboard
Metric: Accuracy (%) on comparative short-answer questions whose answer names one of the two tracks, exact match on the track identifier; the 2,010-pair random subset on which every multi-audio model ran; each model hears both tracks and the question end to end; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o-Audio | 84.4 |
| 2 | Jamendo-MT-QA Qwen2-Audio (checkpoint unspecified) | 77.7 |
| 3 | Qwen3-Omni | 75.5 |
| 4 | GPT-4o-mini-Audio | 73.2 |
Interactive version: theaggregate.ai/benchmark?slug=jamendo-mt-qa-short-answer-questions · How It Works · Data refreshed daily, snapshot 2026-10-07.