Social Gym - Bluffing Games: leaderboard
Metric: Bradley-Terry Elo rating (regularized fit on the pooled cross-model pairwise outcomes of four hidden-state bluffing games (Liar's Dice, Skull, Coup, Sheriff of Nottingham), scale 400, anchored at a mean of 1000 over the seven evaluated models; seats and roles balanced over every model combination; relative to this model pool only). Source: arxiv.org. Saturation forecast: Around October 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 1080 |
| 2 | GPT-4o | 1025 |
| 3 | GPT-4o Mini | 996 |
| 4 | Qwen 3 32B | 982 |
| 5 | Qwen 3 4B | 976 |
| 6 | Gemma 3 27B | 973 |
| 7 | Qwen 2.5 3B Instruct | 964 |
Interactive version: theaggregate.ai/benchmark?slug=social-gym-bluffing-games · How It Works · Data refreshed daily, snapshot 2026-09-29.