TSHA: leaderboard

Metric: Final average (%): mean of the open-ended QA overall (GPT-4o judge rubric: 0.7 accuracy, 0.2 conciseness, 0.1 coherence) and the choice-question overall (mean of yes/no and four-option accuracy) on the 1,707-item TSHA test set (existing indoor datasets, new photos, internet and AIGC images, Sora videos, Hunyuan panoramas); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScoreOverall rank
1Claude 3.7 Sonnet80.6#241
2InternVL3-8B80.1#606
3Claude Sonnet 480#194
4Qwen 2.5 VL 32B Instruct79.9#443
5Qwen 2.5 VL 7B Instruct76.5#643
6Gemini 2.5 Flash61.5#237
7Gemini 2.5 Pro61.4#145
8Gemma 3 27B (IT)39.8#509
9Gemma 3 12B (IT)34.7#655
10Mistral Small 3.133.6#600
11Gemma 3 4B (IT)33.4#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=tsha · How It Works · Data refreshed daily, snapshot 2026-10-11.