HM-Bench (Report): leaderboard

Metric: Average accuracy (%) over all tasks of HM-Bench's 19,337 hyperspectral multiple-choice questions (2,178 samples, 13 task categories) with report-only input: a structured text report of spectral and spatial statistics computed from the cube; zero-shot, one option letter as output, temperature 0, at most 64 new tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 17 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini37.45
2InternVL3-14B37.26
3Claude Sonnet 4.636.57
4InternVL2-8B36.17
5Grok 431.68
6Qwen 2.5 VL 7B Instruct31.03

Interactive version: theaggregate.ai/benchmark?slug=hm-bench-report · How It Works · Data refreshed daily, snapshot 2026-10-07.