OmniToM - Belief Labeling - Content Type: leaderboard
Metric: Belief-labeling accuracy (%): exact-match accuracy of the predicted schema label on the given gold belief propositions, macro-averaged over stories, for the content type (location, physical state, identity, epistemic, desire, emotion, trait or event) dimension, on the 895 OmniToM stories from seven ToMBench categories (22,343 gold belief propositions), zero-shot TELeR level 3 prompts; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 85.97 |
| 2 | Mistral Large 2 (Jul) | 82.83 |
| 3 | GPT-5 | 79.96 |
| 4 | Mistral Small 3 | 76.01 |
| 5 | Gemma 3 27B (IT) | 73.5 |
| 6 | Llama 3.3 70B Instruct | 72.35 |
| 7 | Qwen 3 32B | 71.27 |
| 8 | Qwen 3 8B | 51.43 |
| 9 | Llama 3.1 8B Instruct | 48.63 |
Interactive version: theaggregate.ai/benchmark?slug=omnitom-belief-labeling-content-type · How It Works · Data refreshed daily, snapshot 2026-10-07.