Social Gym - Bargaining: leaderboard
Metric: Bradley-Terry Elo rating (regularized fit on the pooled cross-model pairwise outcomes of the Bargaining game (structured negotiation over a divisible payoff, the one competitive economic game), scale 400, anchored at a mean of 1000 over the seven evaluated models; seats and roles balanced over every model combination; relative to this model pool only). Source: arxiv.org. Saturation forecast: Around November 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 1131 |
| 2 | GPT-4o Mini | 1067 |
| 3 | Qwen 2.5 3B Instruct | 1060 |
| 4 | GPT-5 Mini | 1039 |
| 5 | Qwen 3 4B | 1031 |
| 6 | Qwen 3 32B | 837 |
| 7 | Gemma 3 27B | 831 |
Interactive version: theaggregate.ai/benchmark?slug=social-gym-bargaining · How It Works · Data refreshed daily, snapshot 2026-09-29.