SciHazard (Deep Research Agents) - Harm When Not Refused: leaderboard
Metric: DeHarm-Score over non-refused answers (0-5, lower is safer; the mean over the non-refused answers to 2,400 hazardous questions in 12 scientific disciplines of hazard level 1 to 5 x (executability + net-new risk, each 1 to 4)/8, executability from importance-weighted checklists and net-new risk from retrieval-checked claims, judged by Qwen3.5-35B-A3B; deep-research agents and agent frameworks with web retrieval at provider defaults). Source: arxiv.org. Saturation forecast: Around May 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Sonar Pro | 2.61 |
| 2 | O3 Deep Research | 2.93 |
Interactive version: theaggregate.ai/benchmark?slug=scihazard-deep-research-agents-harm-when-not-refused · How It Works · Data refreshed daily, snapshot 2026-09-29.