PinpointQA: leaderboard

Metric: Average score (0-1 task scores, times 100): arithmetic mean of the four task micro scores (target presence verification, nearest reference identification, fine-grained spatial description, structured spatial prediction) of the PinpointQA evaluation split (2,019 QA pairs from held-out ScanNet++ and ScanNet200 scenes; 64 uniformly sampled RGB frames per indoor video, no depth or 3D input, greedy decoding); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.

Top models

#ModelScore
1Kimi K2.542
2Qwen 3 VL 8B Instruct39
3GPT-5.438

Interactive version: theaggregate.ai/benchmark?slug=pinpointqa · How It Works · Data refreshed daily, snapshot 2026-10-07.