GST-Bench - Egocentric Distance (Semantic): leaderboard

Metric: Mean relative accuracy (%; distance to a previously seen object within relative-error thresholds 0.50 to 0.95, averaged; target named by its category; zero-shot, greedy decoding, official prompt templates, over human-verified questions on 6,790 minutes of simulator-rendered egocentric exploration video). Source: arxiv.org. Saturation forecast: Around August 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro35.4
2GPT-527.39
3Gemini 3 Pro27.39
4Qwen 3 VL 32B16.19
5Seed 1.813.32
6GPT-4o10.13
7InternVL3.5-8B8.36

Interactive version: theaggregate.ai/benchmark?slug=gst-bench-egocentric-distance-semantic · How It Works · Data refreshed daily, snapshot 2026-09-29.