OpenVLM POPE: leaderboard

OpenCompass OpenVLM evaluation of POPE: polling-based object hallucination assessment for vision-language models, testing whether models falsely claim nonexistent objects are present.

Metric: Overall (%). Source: huggingface.co. Status: saturated. 216 models tracked.

Top models

#ModelScore
1InternVL3-78B90.5
2InternVL3-8B90.4
3InternVL3-38B89.2
4Gemini 1.5 Flash88.5
5Qwen 2 VL 7B88.4
6Gemini 1.5 Pro88.2
7Llama 3.2 11B Instruct88.1
8Qwen 2 VL 2B87.3
9Qwen 2 VL 72B87.2
10GPT-4.1 (2025-04-14)86.4
11Llama 3.2 90B Vision Instruct86.3
12GPT-4.1 Mini85.6
13Gemma 3 12B85.5
14Gemma 3 4B84.6
15InternVL2-8B84.2

Interactive version: theaggregate.ai/benchmark?slug=openvlm-pope · How It Works · Data refreshed daily, snapshot 2026-09-05.