GeoDisaster - Answer Accuracy: leaderboard

Metric: Answer accuracy (%; final answer judged correct) on the GeoDisaster test instances (verified operational disaster geo-intelligence questions over optical and SAR imagery, raster masks, vector geometries, road networks and exposure layers, in five task families), each model as a single agent with structured prompting, the shared 18-tool registry and execution budget; a GPT-5.5 judge scores every trajectory; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.561.99
2GPT-560.48
3GPT-4o60.19
4O4 Mini26.64
5Qwen 2.5 7B Instruct2.53
6Qwen 2.5 VL 7B Instruct1.5
7Qwen 3 4B 2507 Instruct1.16
8Llama 3.1 8B Instruct0.25
9Mistral 7B Instruct (v0.3)0.25

Interactive version: theaggregate.ai/benchmark?slug=geodisaster-answer-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.