LLM Public Goods Game — leaderboard

Classic economic cooperation experiment: LLMs decide how many tokens to contribute to a shared pool across 310 multi-player, multi-round games with varying multiplication factors. Measures cooperative vs. self-interested behavior.

Metric: Avg. Contribution (%). Source: github.com. 21 models tracked.

Top models

#ModelScore
1Gemini 2.0 Flash (Preview)45.2
2Claude 3.5 Haiku40.97
3Gemma 2 27B31.16
4Gemini 1.5 Flash30.2
5GPT-4o Mini24.82
6Gemini 1.5 Pro (Sept)24.64
7Qwen 2.5 72B23.59
8Claude 3.5 Sonnet (20241022)20.89
9Gemini 2.0 Flash (01-21) (Thinking)20
10Grok 2 (1212)15.83
11DeepSeek V315.5
12GPT-4o14.7
13Llama 3.1 405B3.16
14O1 Mini2.59
15O12.55

Interactive version: theaggregate.ai/benchmark?slug=llm-public-goods-game · How the rankings work · Data refreshed daily, snapshot 2026-07-22.