HM-Bench (PCA): leaderboard

Metric: Average accuracy (%) over all tasks of HM-Bench's 19,337 hyperspectral multiple-choice questions (2,178 samples, 13 task categories) with PCA-only input: a false-colour image of the cube's leading principal components; zero-shot, one option letter as output, temperature 0, at most 64 new tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 17 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini42.35
2Claude Sonnet 4.639.48
3InternVL2-8B38.09
4InternVL3-14B38.09
5Grok 432.78
6Qwen 2.5 VL 7B Instruct29.36

Interactive version: theaggregate.ai/benchmark?slug=hm-bench-pca · How It Works · Data refreshed daily, snapshot 2026-10-07.