GST-Bench - Global Position (Semantic): leaderboard

Metric: Mean accuracy (%; predicted top-down map position of a previously seen object within 100 to 300 pixels, averaged over five thresholds; target named by its category; zero-shot, greedy decoding, official prompt templates, over human-verified questions on 6,790 minutes of simulator-rendered egocentric exploration video). Source: arxiv.org. Saturation forecast: Around May 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Pro43.67
2GPT-543.23
3Seed 1.836.07
4Gemini 2.5 Pro32.4
5Qwen 3 VL 32B26.55
6GPT-4o25.24
7InternVL3.5-8B14.85

Interactive version: theaggregate.ai/benchmark?slug=gst-bench-global-position-semantic · How It Works · Data refreshed daily, snapshot 2026-09-29.