HypoArena - Safety Investigation: leaderboard

Metric: Rating on the 103 safety investigation cases: Bradley-Terry-Davidson rating (base 1500, open-ended) from reference-independent pairwise judgments by seed-2.0-pro, both presentation orders, with the source-derived reference hypothesis set as an anonymous competitor; baseline single-pass generation; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (High)1721.6
2Claude Opus 4.6 (High)1668.4
3Kimi K2.61633.5
4GPT-5.4 (High)1630.7
5GLM-5.11560
6DeepSeek V4 Pro (High)1551.4
7DeepSeek V4 Flash (High)1489.3
8MiniMax-M2.71453.9
9GLM-51445.2
10GPT-5.4 Mini (High)1437.7
11Gemini 3.1 Pro (Preview) (High)1387.2
12MiniMax-M2.51369.7
13Kimi K2.5 (Thinking)1289.4
14Gemini 3 Flash (High)1262.1

Interactive version: theaggregate.ai/benchmark?slug=hypoarena-safety-investigation · How It Works · Data refreshed daily, snapshot 2026-09-29.