GeoDisaster: leaderboard

Metric: Task success rate (%) on the GeoDisaster test instances (verified operational disaster geo-intelligence questions over optical and SAR imagery, raster masks, vector geometries, road networks and exposure layers, in five task families), each model as a single agent with structured prompting, the shared 18-tool registry and execution budget; a GPT-5.5 judge scores every trajectory; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.555.06
2GPT-546.27
3GPT-4o42.65
4O4 Mini29.12
5Qwen 2.5 7B Instruct3.11
6Qwen 3 4B 2507 Instruct1.16
7Qwen 2.5 VL 7B Instruct0.95
8Llama 3.1 8B Instruct0.15
9Mistral 7B Instruct (v0.3)0.15

Interactive version: theaggregate.ai/benchmark?slug=geodisaster · How It Works · Data refreshed daily, snapshot 2026-09-29.