Uruqi: leaderboard
Metric: Overall accuracy (%): mean of the three stage scores (self-motion, object mapping, state operations), subclasses weighted equally within each stage, over 52k questions on 2.7k simulated camera-trajectory episodes; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-6 Astra | 50.08 |
| 2 | Gemini 3.1 Pro (Preview) | 27.37 |
| 3 | GPT-5.5 | 24.5 |
| 4 | Qwen 3.8 Max | 23.58 |
| 5 | Qwen 3.8 27B | 16.94 |
| 6 | InternVL3-8B | 15.84 |
Interactive version: theaggregate.ai/benchmark?slug=uruqi · How It Works · Data refreshed daily, snapshot 2026-10-02.