OpenVLM POPE — leaderboard

OpenCompass OpenVLM evaluation of POPE: polling-based object hallucination assessment for vision-language models, testing whether models falsely claim nonexistent objects are present.

Metric: Overall (%). Source: huggingface.co. Status: saturated. 216 models tracked.

Top models

#ModelScore
1InternVL3-78B90.5
2InternVL3-8B90.4
3InternVL3-14B89.4
4InternVL3-38B89.2
5Ovis2-8B88.6
6Gemini 1.5 Flash88.5
7Qwen 2 VL 7B88.4
8Gemini 1.5 Pro88.2
9Llama 3.2 11B Instruct88.1
10Qwen 2 VL 2B87.3
11Qwen 2 VL 72B87.2
12GPT-4.1 (2025-04-14)86.4
13Llama 3.2 90B Vision Instruct86.3
14GPT-4.1 Mini85.6
15Gemma 3 12B85.5

Interactive version: theaggregate.ai/benchmark?slug=openvlm-pope · How the rankings work · Data refreshed daily, snapshot 2026-07-22.