GPSBench (Applied) - Relative Position: leaderboard

Metric: Score (%) on the relative position (compass direction between two named cities of one continent, 200 to 3,000 km apart) task of GPSBench's Applied Track (geographic reasoning that combines coordinates or place names with world knowledge); 1,020 test samples; zero-shot from the model's own knowledge, no tools, no chain-of-thought or few-shot prompting, temperature 0; choice tasks score accuracy and numerical tasks one minus the mean absolute percentage error; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash (Non-reasoning)87.2#237 (Gemini 2.5 Flash)
2GPT-4.164.8#240
3Gemini 2.5 Pro62.2#145
4GPT-5.1 (Low)58.5#131 (GPT-5.1)
5GPT-5 Mini (Low)56.6#176 (GPT-5 Mini)
6GPT-4.1 Mini53.1#346
7GPT-5 Nano (Low)52.6#415 (GPT-5 Nano)
8Claude Haiku 4.551.7#271
9Qwen 3 235B A22B 2507 Instruct50#291
10Mistral Large 347.5#388
11Mistral Small 344.8#616
12Qwen 3 30B A3B 2507 Instruct42.5#464
13Qwen 3 14B (Non-reasoning)38.2#524 (Qwen 3 14B)
14Qwen 3 8B (Non-reasoning)35.1#667 (Qwen 3 8B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gpsbench-applied-relative-position · How It Works · Data refreshed daily, snapshot 2026-10-11.