SciHazard (Deep Research Agents) - Harm When Not Refused: leaderboard

Metric: DeHarm-Score over non-refused answers (0-5, lower is safer; the mean over the non-refused answers to 2,400 hazardous questions in 12 scientific disciplines of hazard level 1 to 5 x (executability + net-new risk, each 1 to 4)/8, executability from importance-weighted checklists and net-new risk from retrieval-checked claims, judged by Qwen3.5-35B-A3B; deep-research agents and agent frameworks with web retrieval at provider defaults). Source: arxiv.org. Saturation forecast: Around May 2027. 9 models tracked.

Top models

#ModelScore
1Sonar Pro2.61
2O3 Deep Research2.93

Interactive version: theaggregate.ai/benchmark?slug=scihazard-deep-research-agents-harm-when-not-refused · How It Works · Data refreshed daily, snapshot 2026-09-29.