BehaviorBench - Cross-Game Prediction (Distribution): leaderboard

Metric: Given the first-round moves of each subject in other games, predict the first-round moves in a target game, on MobLab economic-game records of human players (dictator, ultimatum, trust, public goods, bomb risk, beauty contest and push/pull games): Wasserstein distance (0-100) between the model-predicted and the observed human choice distributions, with choices normalized to 0-100 and averaged over games; lower is better. Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)10.3
2DeepSeek V3.212.3
3GPT-5.4 (High)14
4Llama 3.3 70B Instruct14.9
5GPT-4.116.3
6Claude Haiku 4.517.3
7Gemini 3.1 Flash Lite (Preview)17.3
8GPT-5.4 Mini (High)18
9Claude Sonnet 4.619.3
10Claude Opus 4.620.1
11Qwen 3 4B20.1

Interactive version: theaggregate.ai/benchmark?slug=behaviorbench-cross-game-prediction-distribution · How It Works · Data refreshed daily, snapshot 2026-09-29.