UrbanWell - Land Use Classification: leaderboard
Metric: Accuracy (%) of choosing the land-use category of a grid cell from same-year satellite and street-view images (multiple choice), cells sampled evenly across 38 mostly European cities, zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.0 Flash | 66 |
| 2 | Gemma 3 4B | 60 |
| 3 | Gemma 3 27B | 59 |
| 4 | GPT-5 Nano | 58 |
| 5 | Gemma 3 12B | 55 |
| 6 | Qwen 2.5 VL 7B Instruct | 51 |
| 7 | Nova Lite (v1) | 49 |
| 8 | Qwen 2.5 VL 32B Instruct | 48 |
| 9 | Phi-4 Multimodal Instruct | 25 |
| 10 | Llama 4 Scout | 7 |
Interactive version: theaggregate.ai/benchmark?slug=urbanwell-land-use-classification · How It Works · Data refreshed daily, snapshot 2026-09-29.