GeoDisaster: leaderboard
Metric: Task success rate (%) on the GeoDisaster test instances (verified operational disaster geo-intelligence questions over optical and SAR imagery, raster masks, vector geometries, road networks and exposure layers, in five task families), each model as a single agent with structured prompting, the shared 18-tool registry and execution budget; a GPT-5.5 judge scores every trajectory; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 55.06 |
| 2 | GPT-5 | 46.27 |
| 3 | GPT-4o | 42.65 |
| 4 | O4 Mini | 29.12 |
| 5 | Qwen 2.5 7B Instruct | 3.11 |
| 6 | Qwen 3 4B 2507 Instruct | 1.16 |
| 7 | Qwen 2.5 VL 7B Instruct | 0.95 |
| 8 | Llama 3.1 8B Instruct | 0.15 |
| 9 | Mistral 7B Instruct (v0.3) | 0.15 |
Interactive version: theaggregate.ai/benchmark?slug=geodisaster · How It Works · Data refreshed daily, snapshot 2026-09-29.