OmniToM - Belief Labeling - Content Type: leaderboard

Metric: Belief-labeling accuracy (%): exact-match accuracy of the predicted schema label on the given gold belief propositions, macro-averaged over stories, for the content type (location, physical state, identity, epistemic, desire, emotion, trait or event) dimension, on the 895 OmniToM stories from seven ToMBench categories (22,343 gold belief propositions), zero-shot TELeR level 3 prompts; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash85.97
2Mistral Large 2 (Jul)82.83
3GPT-579.96
4Mistral Small 376.01
5Gemma 3 27B (IT)73.5
6Llama 3.3 70B Instruct72.35
7Qwen 3 32B71.27
8Qwen 3 8B51.43
9Llama 3.1 8B Instruct48.63

Interactive version: theaggregate.ai/benchmark?slug=omnitom-belief-labeling-content-type · How It Works · Data refreshed daily, snapshot 2026-10-07.