HM-Bench (VSR2): leaderboard

Metric: Average accuracy (%) over all tasks of HM-Bench's 19,337 hyperspectral multiple-choice questions (2,178 samples, 13 task categories) with full VSR2 input: the RGB, PCA and report views together with the authors' fixed cross-view fusion prompt; zero-shot, one option letter as output, temperature 0, at most 64 new tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 17 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini44.28
2Claude Sonnet 4.644.01
3InternVL3-14B40.05
4Grok 439.5
5Qwen 2.5 VL 7B Instruct39.41
6InternVL2-8B39.12

Interactive version: theaggregate.ai/benchmark?slug=hm-bench-vsr2 · How It Works · Data refreshed daily, snapshot 2026-10-07.