HypoArena - Social Science: leaderboard

Metric: Rating on the 163 social science cases: Bradley-Terry-Davidson rating (base 1500, open-ended) from reference-independent pairwise judgments by seed-2.0-pro, both presentation orders, with the source-derived reference hypothesis set as an anonymous competitor; baseline single-pass generation; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (High)1671.5
2GPT-5.4 (High)1602.3
3GLM-5.11591.6
4Claude Sonnet 4.6 (High)1582.7
5DeepSeek V4 Pro (High)1564.8
6Kimi K2.61559.6
7DeepSeek V4 Flash (High)1513.8
8MiniMax-M2.71492.5
9GLM-51452.6
10MiniMax-M2.51427.4
11Gemini 3 Flash (High)1418.2
12Gemini 3.1 Pro (Preview) (High)1415.7
13GPT-5.4 Mini (High)1393.7
14Kimi K2.5 (Thinking)1322.7

Interactive version: theaggregate.ai/benchmark?slug=hypoarena-social-science · How It Works · Data refreshed daily, snapshot 2026-09-29.