EnactToM (Hard) - Literal ToM Probes: leaderboard

Metric: Literal ToM probe accuracy (%, Avg: mean per-run success over three runs of each task): at the end of each embodied multi-agent episode every agent is asked explicit questions about what another agent knows, derived from the task's epistemic goal operators, over the overall scope (22 cooperative and 18 mixed-motive tasks) of the matched EnactToM hard subset (seed-task failure ratio 0.9); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.

Top models

#ModelScore
1O352.5
2GPT-5.444.2
3Kimi K2.544.2
4DeepSeek V3.236.7
5GPT-5.4 Mini31.7

Interactive version: theaggregate.ai/benchmark?slug=enacttom-hard-literal-tom-probes · How It Works · Data refreshed daily, snapshot 2026-10-07.