SciHazard - Harm When Not Refused: leaderboard

Metric: DeHarm-Score over non-refused answers (0-5, lower is safer; the mean over the non-refused answers to 2,400 hazardous questions in 12 scientific disciplines of hazard level 1 to 5 x (executability + net-new risk, each 1 to 4)/8, executability from importance-weighted checklists and net-new risk from retrieval-checked claims, judged by Qwen3.5-35B-A3B; standard LLMs at temperature 0). Source: arxiv.org. Saturation forecast: Around 2032. 19 models tracked.

Top models

#ModelScore
1GPT-5.41.78
2GPT-5.21.91
3Claude Opus 4.61.95
4Claude Sonnet 4.61.98
5Llama 4 Scout2.07
6Llama 3.1 70B Instruct2.21
7Gemini 3.1 Pro (Preview)2.33
8Gemini 3 Flash2.45
9Claude Haiku 4.52.45
10GLM-5.12.52
11Qwen 3.5 397B A17B2.64
12O3 Mini2.69
13Kimi K2.52.75
14Qwen 3.5 35B A3B2.8
15DeepSeek R12.9

Interactive version: theaggregate.ai/benchmark?slug=scihazard-harm-when-not-refused · How It Works · Data refreshed daily, snapshot 2026-09-29.