TSHA - Choice Questions: leaderboard
Metric: Choice-question overall (%): mean of yes/no hazard-presence accuracy and four-option single-choice accuracy on the 1,707-item TSHA test set (existing indoor datasets, new photos, internet and AIGC images, Sora videos, Hunyuan panoramas); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | InternVL3-8B | 91.3 | #606 |
| 2 | Claude 3.7 Sonnet | 87.1 | #241 |
| 3 | Qwen 2.5 VL 32B Instruct | 84 | #443 |
| 4 | Qwen 2.5 VL 7B Instruct | 81.8 | #643 |
| 5 | Claude Sonnet 4 | 81.2 | #194 |
| 6 | Mistral Small 3.1 | 47.8 | #600 |
| 7 | Gemini 2.5 Flash | 45.8 | #237 |
| 8 | Gemini 2.5 Pro | 44.4 | #145 |
| 9 | Gemma 3 27B (IT) | 41.8 | #509 |
| 10 | Gemma 3 12B (IT) | 33 | #655 |
| 11 | Gemma 3 4B (IT) | 30.2 | #971 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=tsha-choice-questions · How It Works · Data refreshed daily, snapshot 2026-10-11.