LLM Public Goods Game — leaderboard
Classic economic cooperation experiment: LLMs decide how many tokens to contribute to a shared pool across 310 multi-player, multi-round games with varying multiplication factors. Measures cooperative vs. self-interested behavior.
Metric: Avg. Contribution (%). Source: github.com. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.0 Flash (Preview) | 45.2 |
| 2 | Claude 3.5 Haiku | 40.97 |
| 3 | Gemma 2 27B | 31.16 |
| 4 | Gemini 1.5 Flash | 30.2 |
| 5 | GPT-4o Mini | 24.82 |
| 6 | Gemini 1.5 Pro (Sept) | 24.64 |
| 7 | Qwen 2.5 72B | 23.59 |
| 8 | Claude 3.5 Sonnet (20241022) | 20.89 |
| 9 | Gemini 2.0 Flash (01-21) (Thinking) | 20 |
| 10 | Grok 2 (1212) | 15.83 |
| 11 | DeepSeek V3 | 15.5 |
| 12 | GPT-4o | 14.7 |
| 13 | Llama 3.1 405B | 3.16 |
| 14 | O1 Mini | 2.59 |
| 15 | O1 | 2.55 |
Interactive version: theaggregate.ai/benchmark?slug=llm-public-goods-game · How the rankings work · Data refreshed daily, snapshot 2026-07-22.