ThermEval - Thermal Reasoning (Two People): leaderboard

Metric: Accuracy (times 100, 0-100) on Task 5, whether the chest, forehead or nose of the left or the right person is hotter, from a ThermEval-D image with a colorbar (155 questions); zero-shot with fixed one-word or one-number prompts, a single forward pass; Gemini 2.5 models parse free-form answers into the label or number only when the output deviates from the format; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 24 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro74#77
2Llama 3.2 11B Instruct61#1112
3Gemini 2.5 Pro57#145
4Gemini 3 Flash55#93
5InternVL3-14B53#494
6Gemini 2.5 Flash52#237
7InternVL3-8B48#606
8InternVL3-38B48#395
9Qwen 2.5 VL 7B Instruct44#643
10MiniCPM-V-2.641#825
11Qwen 2 VL 7B Instruct41#816
12Qwen 2.5 VL 32B Instruct37#443

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=thermeval-thermal-reasoning-two-people · How It Works · Data refreshed daily, snapshot 2026-10-11.