MuseBench - Stage Performing Arts (Single-Select): leaderboard

Metric: Stage-performing-arts chance-adjusted accuracy (%): raw single-select accuracy minus the per-item chance baseline, rescaled so random guessing scores 0 and a correct answer 1; a negative value is worse than chance; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 28 models tracked.

Top models

#ModelScore
1Claude Opus 4.662.65
2Qwen 3.5 Plus60.69
3InternVL3-78B57.25
4Qwen 3.5 397B A17B56.6
5GPT-5.456.43
6Gemini 3.1 Pro (Preview)49.5
7Qwen2.5-Omni-7B34.43
8Gemma 4 E4B31.44
9InternVL3-8B30.48
10Kimi K2.523.62
11Grok 4.115.9
12GLM-4.5V13.6

Interactive version: theaggregate.ai/benchmark?slug=musebench-stage-performing-arts-single-select · How It Works · Data refreshed daily, snapshot 2026-09-29.