BehaviorBench - Cross-Game Prediction: leaderboard

Metric: Given the first-round moves of one subject in other games, predict the first-round move in a target game, on MobLab economic-game records of human players (dictator, ultimatum, trust, public goods, bomb risk, beauty contest and push/pull games): Mean absolute error of the predicted move in the game choice units; lower is better. Source: arxiv.org. Saturation forecast: Around August 2027. 18 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.625.3
2Claude Opus 4.625.9
3Claude Haiku 4.526
4GPT-5.4 (High)26.1
5Llama 3.3 70B Instruct26.4
6Gemini 3.1 Pro (Preview)26.7
7GPT-4.127.1
8GPT-5.4 Mini (High)27.4
9DeepSeek V3.227.5
10Qwen 3 4B27.9
11Gemini 3.1 Flash Lite (Preview)29.4

Interactive version: theaggregate.ai/benchmark?slug=behaviorbench-cross-game-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.