CoopEval - Reputation (First-Order): leaderboard

Metric: Normalized mean payoff (open scale; below 0 when exploited, above 1 possible under side payments) under first-order reputation (random re-matching each round with co-players' own past 3 rounds visible), aggregated over four social dilemmas (Prisoner's Dilemma, Public Goods, Traveler's Dilemma, Trust Game) after shifting and rescaling payoffs so that 0 is everyone defecting and 1 is everyone playing the most cooperative action; the model's mean payoff across all cross-play match-ups with the six tested LLM agents (a uniform population), each instructed to maximize its own points, temperature 1, three repeats per combination; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 6 models tracked.

Top models

#ModelScore
1Qwen 3 30B A3B 2507 Instruct0.4
2GPT-4o (2024-05-13)0.34
3GPT-5.2 (Low)0.33
4Gemini 3 Flash (Medium)0.28

Interactive version: theaggregate.ai/benchmark?slug=coopeval-reputation-first-order · How It Works · Data refreshed daily, snapshot 2026-10-07.