Werewolf Benchmark — leaderboard

ELO-rated arena where LLMs play the hidden-role social deduction game Werewolf, testing deception, persuasion, and reasoning about other players' hidden identities.

Metric: ELO. Source: werewolf.foaster.ai. Status: saturation imminent. 10 models tracked.

Top models

#ModelScore
1GPT-51529
2Gemini 2.5 Pro1243
3Grok 4 Fast (Reasoning)1231
4Gemini 2.5 Flash1222
5Kimi K2 09051189
6Grok 41178
7Qwen 3 235B A22B 2507 Instruct1150
8Kimi K21133
9GPT-5 Mini1120
10GPT-OSS-120B971

Interactive version: theaggregate.ai/benchmark?slug=werewolf-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.