UniVerseBench (Thinking Mode): leaderboard
Metric: Strict accuracy (%; multiple-choice questions on folk and traditional music recordings from 372 recordings in over 38 languages, 5,042 questions in total, verified by music experts; thinking mode, a valid reasoning chain and the correct final answer both required). Source: arxiv.org. Saturation forecast: Around December 2026. 2 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3 Omni 30B A3B Instruct | 47.5 |
| 2 | Qwen2.5-Omni-7B | 33.9 |
Interactive version: theaggregate.ai/benchmark?slug=universebench-thinking-mode · How It Works · Data refreshed daily, snapshot 2026-09-26.