TSHA - Open-Ended Hazard QA: leaderboard

Metric: Open-ended QA overall (%): GPT-4o rubric score, 0.7 hazard accuracy + 0.2 conciseness + 0.1 coherence, on the 1,707-item TSHA test set (existing indoor datasets, new photos, internet and AIGC images, Sora videos, Hunyuan panoramas); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 478.8#194
2Gemini 2.5 Pro78.3#145
3Gemini 2.5 Flash77.1#237
4Qwen 2.5 VL 32B Instruct75.7#443
5Claude 3.7 Sonnet74#241
6Qwen 2.5 VL 7B Instruct71.1#643
7InternVL3-8B68.9#606
8Gemma 3 27B (IT)37.7#509
9Gemma 3 4B (IT)36.7#971
10Gemma 3 12B (IT)36.4#655
11Mistral Small 3.119.4#600

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=tsha-open-ended-hazard-qa · How It Works · Data refreshed daily, snapshot 2026-10-11.