GeoDisaster - Answer Accuracy: leaderboard
Metric: Answer accuracy (%; final answer judged correct) on the GeoDisaster test instances (verified operational disaster geo-intelligence questions over optical and SAR imagery, raster masks, vector geometries, road networks and exposure layers, in five task families), each model as a single agent with structured prompting, the shared 18-tool registry and execution budget; a GPT-5.5 judge scores every trajectory; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 61.99 |
| 2 | GPT-5 | 60.48 |
| 3 | GPT-4o | 60.19 |
| 4 | O4 Mini | 26.64 |
| 5 | Qwen 2.5 7B Instruct | 2.53 |
| 6 | Qwen 2.5 VL 7B Instruct | 1.5 |
| 7 | Qwen 3 4B 2507 Instruct | 1.16 |
| 8 | Llama 3.1 8B Instruct | 0.25 |
| 9 | Mistral 7B Instruct (v0.3) | 0.25 |
Interactive version: theaggregate.ai/benchmark?slug=geodisaster-answer-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.