SelfCheckGPT: leaderboard

Black-box hallucination detection benchmark/method using repeated samples from the same model. Inconsistent sampled claims are treated as signals of possible factual fabrication.

Metric: Resistance (100 - SelfCheckGPT). Source: arxiv.org. 9 models tracked.

Top models

#ModelScore
1InstructBLIP88.32
2LLaVA-v1.587.43
3FAVOR86.22
4Llama-VID85.14
5MPLUG-Owl283.93
6Chat-Univi83.2
7Video-LLaVA80.03
8Video-Llama77.01
9Valley74.31

Interactive version: theaggregate.ai/benchmark?slug=selfcheckgpt · How It Works · Data refreshed daily, snapshot 2026-09-05.