Sens-VisualNews - Strict: leaderboard

Metric: Top-1 accuracy (%) on the test split of the strict subset (sensational images with unanimous annotator agreement and as many non-sensational images), zero-shot yes or no answer to whether the news image is sensational, greedy decoding, each model with the prompt and prompt format that scored best on the development split; balanced classes, chance 50; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 12 models tracked.

Top models

#ModelScore
1LLaVA-OneVision-1.5-8B93.9
2Qwen3-VL-8B92.8
3Qwen3-VL-2B92.7
4InternVL3.5-8B92.1
5LLaVA-OneVision-7B91.8
6Qwen3-VL-4B90.8
7LLaVA-OneVision-0.5B90.8
8LLaVA-OneVision-1.5-4B89.7
9InternVL3.5-4B88.8
10InternVL3.5-2B86.2

Interactive version: theaggregate.ai/benchmark?slug=sens-visualnews-strict · How It Works · Data refreshed daily, snapshot 2026-10-07.