CoopEval - Contracting: leaderboard
Metric: Normalized mean payoff (open scale; below 0 when exploited, above 1 possible under side payments) under contracting (agents propose outcome-conditional payment contracts, approve one by vote and play under it), aggregated over four social dilemmas (Prisoner's Dilemma, Public Goods, Traveler's Dilemma, Trust Game) after shifting and rescaling payoffs so that 0 is everyone defecting and 1 is everyone playing the most cooperative action; the model's mean payoff across all cross-play match-ups with the six tested LLM agents (a uniform population), each instructed to maximize its own points, temperature 1, three repeats per combination; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Medium) | 1.06 |
| 2 | GPT-5.2 (Low) | 0.83 |
| 3 | Qwen 3 30B A3B 2507 Instruct | 0.78 |
| 4 | GPT-4o (2024-05-13) | 0.45 |
Interactive version: theaggregate.ai/benchmark?slug=coopeval-contracting · How It Works · Data refreshed daily, snapshot 2026-10-07.