GeoBrowse - Level 1 State (Direct): leaderboard

Metric: Accuracy (%) on the state-level targets of GeoBrowse Level 1 (199 street-scene instances whose location must be composed from weak, fragmented visual cues), judged correct by an LLM judge with the Humanity's Last Exam judging prompt, mean of 3 runs, under direct inference: the model answers in a single pass without tool calls; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 12 models tracked.

Top models

#ModelScore
1Gemini 3 Pro27.3
2Claude Opus 4.523.6
3Gemini 2.5 Pro21.8
4GPT-521.8
5Gemini 2.5 Flash18.2
6GPT-4o16.4
7GPT-4.114.5
8Llama 3.2 90B10.9
9GPT-5.29.1
10Qwen 2.5 VL 7B6.3

Interactive version: theaggregate.ai/benchmark?slug=geobrowse-level-1-state-direct · How It Works · Data refreshed daily, snapshot 2026-10-07.