MoHallBench - Semantic Similarity Hallucination: leaderboard

Metric: Q-PairAcc (%; share of adversarial pairs in which the model answers both the positive and the negative (reversed) yes/no query correctly; 32 uniformly sampled frames, temperature 0; semantically related but physically distinct actions). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Qwen 3 VL 8B Instruct39.3
2Molmo2-8B27
3InternVL3-8B25.8
4Qwen 2.5 VL 7B Instruct20.7
5Qwen 2 VL 7B Instruct16.4

Interactive version: theaggregate.ai/benchmark?slug=mohallbench-semantic-similarity-hallucination · How It Works · Data refreshed daily, snapshot 2026-09-29.