Obshazard-bench - Multi-faceted Impact Quantification: leaderboard
Metric: Normalized task score (%; x100 of the 0-1 task-aware soft score following the Earth AI protocol: categorical answers by semantic or lexical consistency, numeric and interval answers by distance to the ground truth; equal-weight aggregation across timesteps, subtasks, lifecycle stages and disaster categories; zero-shot, uniform prompts; raw satellite sounding streams with ground-station, disaster-record and socio-economic inputs; post-disaster magnitude deduction, humanitarian burden and socio-economic intensity estimation, 2,465 samples). Source: arxiv.org. Saturation forecast: Around August 2028. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 29.77 |
| 2 | Claude Opus 4.8 | 22.97 |
| 3 | Kimi K2.6 | 22.75 |
| 4 | Qwen 3.5 397B A17B | 15.91 |
Interactive version: theaggregate.ai/benchmark?slug=obshazard-bench-multi-faceted-impact-quantification · How It Works · Data refreshed daily, snapshot 2026-09-26.