Social Gym - Hidden-Role Deduction: leaderboard
Metric: Bradley-Terry Elo rating (regularized fit on the pooled cross-model pairwise outcomes of six hidden-role deduction games (Chameleon, Insider, Spyfall, Undercover, Resistance, Werewolves), scale 400, anchored at a mean of 1000 over the seven evaluated models; seats and roles balanced over every model combination; relative to this model pool only). Source: arxiv.org. Saturation forecast: Around September 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 1132 |
| 2 | GPT-4o | 1070 |
| 3 | Gemma 3 27B | 1032 |
| 4 | GPT-4o Mini | 998 |
| 5 | Qwen 3 32B | 949 |
| 6 | Qwen 3 4B | 935 |
| 7 | Qwen 2.5 3B Instruct | 880 |
Interactive version: theaggregate.ai/benchmark?slug=social-gym-hidden-role-deduction · How It Works · Data refreshed daily, snapshot 2026-09-29.