HM-Bench (RGB): leaderboard

Metric: Average accuracy (%) over all tasks of HM-Bench's 19,337 hyperspectral multiple-choice questions (2,178 samples, 13 task categories) with RGB-only input: a pseudo-RGB rendering of the hyperspectral cube; zero-shot, one option letter as output, temperature 0, at most 64 new tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 17 models tracked.

Top models

#ModelScore
1InternVL3-14B42.48
2Claude Sonnet 4.640.95
3GPT-5.4 Mini39.88
4Qwen 2.5 VL 7B Instruct38.19
5InternVL2-8B36.63
6Grok 434.67

Interactive version: theaggregate.ai/benchmark?slug=hm-bench-rgb · How It Works · Data refreshed daily, snapshot 2026-10-07.