EnactToM (Hard) - Literal ToM Probes: leaderboard
Metric: Literal ToM probe accuracy (%, Avg: mean per-run success over three runs of each task): at the end of each embodied multi-agent episode every agent is asked explicit questions about what another agent knows, derived from the task's epistemic goal operators, over the overall scope (22 cooperative and 18 mixed-motive tasks) of the matched EnactToM hard subset (seed-task failure ratio 0.9); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 | 52.5 |
| 2 | GPT-5.4 | 44.2 |
| 3 | Kimi K2.5 | 44.2 |
| 4 | DeepSeek V3.2 | 36.7 |
| 5 | GPT-5.4 Mini | 31.7 |
Interactive version: theaggregate.ai/benchmark?slug=enacttom-hard-literal-tom-probes · How It Works · Data refreshed daily, snapshot 2026-10-07.