PinpointQA: leaderboard
Metric: Average score (0-1 task scores, times 100): arithmetic mean of the four task micro scores (target presence verification, nearest reference identification, fine-grained spatial description, structured spatial prediction) of the PinpointQA evaluation split (2,019 QA pairs from held-out ScanNet++ and ScanNet200 scenes; 64 uniformly sampled RGB frames per indoor video, no depth or 3D input, greedy decoding); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2.5 | 42 |
| 2 | Qwen 3 VL 8B Instruct | 39 |
| 3 | GPT-5.4 | 38 |
Interactive version: theaggregate.ai/benchmark?slug=pinpointqa · How It Works · Data refreshed daily, snapshot 2026-10-07.