DisasterBench (UAV) - Cascading Risk Reasoning: leaderboard
Metric: Accuracy (%) on cascading risk reasoning (CRR) questions on the 3,000-question test split of low-altitude UAV disaster images (14 disaster scene types), multiple choice scored by strict exact match on the final option letter; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.0 Flash | 81.45 |
| 2 | Gemini 2.5 Pro | 80.32 |
| 3 | GPT-5 | 80.09 |
| 4 | Claude Sonnet 4.5 | 76.02 |
| 5 | InternVL2-8B | 72.4 |
| 6 | GPT-4o | 71.95 |
| 7 | Gemini 2.5 Flash | 71.04 |
| 8 | Qwen 2.5 VL 7B Instruct | 70.14 |
| 9 | Grok 4 | 69.68 |
| 10 | GPT-4o Mini | 62.44 |
| 11 | InternVL3.5-8B | 51.36 |
Interactive version: theaggregate.ai/benchmark?slug=disasterbench-uav-cascading-risk-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.