GPSBench (Applied) - Relative Position: leaderboard
Metric: Score (%) on the relative position (compass direction between two named cities of one continent, 200 to 3,000 km apart) task of GPSBench's Applied Track (geographic reasoning that combines coordinates or place names with world knowledge); 1,020 test samples; zero-shot from the model's own knowledge, no tools, no chain-of-thought or few-shot prompting, temperature 0; choice tasks score accuracy and numerical tasks one minus the mean absolute percentage error; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Flash (Non-reasoning) | 87.2 | #237 (Gemini 2.5 Flash) |
| 2 | GPT-4.1 | 64.8 | #240 |
| 3 | Gemini 2.5 Pro | 62.2 | #145 |
| 4 | GPT-5.1 (Low) | 58.5 | #131 (GPT-5.1) |
| 5 | GPT-5 Mini (Low) | 56.6 | #176 (GPT-5 Mini) |
| 6 | GPT-4.1 Mini | 53.1 | #346 |
| 7 | GPT-5 Nano (Low) | 52.6 | #415 (GPT-5 Nano) |
| 8 | Claude Haiku 4.5 | 51.7 | #271 |
| 9 | Qwen 3 235B A22B 2507 Instruct | 50 | #291 |
| 10 | Mistral Large 3 | 47.5 | #388 |
| 11 | Mistral Small 3 | 44.8 | #616 |
| 12 | Qwen 3 30B A3B 2507 Instruct | 42.5 | #464 |
| 13 | Qwen 3 14B (Non-reasoning) | 38.2 | #524 (Qwen 3 14B) |
| 14 | Qwen 3 8B (Non-reasoning) | 35.1 | #667 (Qwen 3 8B) |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=gpsbench-applied-relative-position · How It Works · Data refreshed daily, snapshot 2026-10-11.