ERGeoBench - Foundational Perception: leaderboard
Metric: Accuracy (%) on the foundational perception diagnostic questions: multiple-choice and true/false questions on architecture, infrastructure, signage, terrain and vegetation in a declared field of view, averaged over the five categories; unified structured prompt with egocentric observations, action history and in-context examples, temperature 0.1, first valid JSON object parsed (invalid outputs count as failures); higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 82.78 |
| 2 | InternVL3-8B | 75.99 |
| 3 | Gemini 2.5 Pro | 70.26 |
| 4 | Gemini 2.0 Flash | 69.08 |
| 5 | Gemini 3 Flash | 66.37 |
| 6 | Qwen 2.5 VL 7B | 61.52 |
Interactive version: theaggregate.ai/benchmark?slug=ergeobench-foundational-perception · How It Works · Data refreshed daily, snapshot 2026-10-07.