Jamendo-MT-QA - Short-Answer Questions: leaderboard

Metric: Accuracy (%) on comparative short-answer questions whose answer names one of the two tracks, exact match on the track identifier; the 2,010-pair random subset on which every multi-audio model ran; each model hears both tracks and the question end to end; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1GPT-4o-Audio84.4
2Jamendo-MT-QA Qwen2-Audio (checkpoint unspecified)77.7
3Qwen3-Omni75.5
4GPT-4o-mini-Audio73.2

Interactive version: theaggregate.ai/benchmark?slug=jamendo-mt-qa-short-answer-questions · How It Works · Data refreshed daily, snapshot 2026-10-07.