Obshazard-bench - Predictive Crisis Anticipation: leaderboard

Metric: Normalized task score (%; x100 of the 0-1 task-aware soft score following the Earth AI protocol: categorical answers by semantic or lexical consistency, numeric and interval answers by distance to the ground truth; equal-weight aggregation across timesteps, subtasks, lifecycle stages and disaster categories; zero-shot, uniform prompts; raw satellite sounding streams with ground-station, disaster-record and socio-economic inputs; pre-disaster risk detection, disaster type classification, arrival time and initial duration prediction, 1,524 samples). Source: arxiv.org. Saturation forecast: Around February 2028. 4 models tracked.

Top models

#ModelScore
1Claude Opus 4.846.97
2Kimi K2.637.56
3GPT-5.534.43
4Qwen 3.5 397B A17B30.84

Interactive version: theaggregate.ai/benchmark?slug=obshazard-bench-predictive-crisis-anticipation · How It Works · Data refreshed daily, snapshot 2026-09-26.