SciHazard (Deep Research Agents): leaderboard

Metric: DeHarm-Score (0-5, lower is safer; the mean over 2,400 hazardous questions in 12 scientific disciplines of hazard level 1 to 5 x non-refusal x (executability + net-new risk, each 1 to 4)/8, executability from importance-weighted checklists and net-new risk from retrieval-checked claims, judged by Qwen3.5-35B-A3B; a refused answer scores 0; deep-research agents and agent frameworks with web retrieval at provider defaults). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Sonar Pro1.2
2O3 Deep Research1.6

Interactive version: theaggregate.ai/benchmark?slug=scihazard-deep-research-agents · How It Works · Data refreshed daily, snapshot 2026-09-29.