Pause-and-Think-B - Scene Understanding: leaderboard

Metric: Binary validity accuracy (%) on the 100 scene-understanding items (object, attribute and action-state queries), GPT-5.1 judge, three runs; higher is better. Source: arxiv.org. Saturation forecast: Around September 2028. 11 models tracked.

Top models

#ModelScore
1GPT-5.255.13
2Qwen 3 VL 235B A22B Instruct53.03
3Gemini 2.5 Pro49.05
4Qwen 3 VL 4B Instruct46.86
5GPT-4o42.58

Interactive version: theaggregate.ai/benchmark?slug=pause-and-think-b-scene-understanding · How It Works · Data refreshed daily, snapshot 2026-09-29.