TimeSpot - Season: leaderboard

Metric: Accuracy (%) of the predicted season on TimeSpot's 1,455 ground-level photographs from 80 countries, each answered zero-shot in a structured multi-field schema from the image alone; a Gemini-2.5-Flash judge normalizes each field (synonyms, abbreviations) and marks it correct or not; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 27 models tracked.

Top models

#ModelScoreOverall rank
1O4 Mini65.81#172
2Qwen 2.5 VL 7B Instruct61.46#643
3GPT-5 Mini58.43#176
4GLM-4.1V-9B (Thinking)58.02#457 (GLM-4.1V-9B)
5GLM-4.5V57.55#339
6Gemini 2.5 Flash (Thinking)51.13#237 (Gemini 2.5 Flash)
7Gemini 2.5 Flash (Non-reasoning)50.92#237 (Gemini 2.5 Flash)
8Gemini 2.0 Flash49.76#331
9Gemini 2.0 Flash (Thinking)49.28#331 (Gemini 2.0 Flash)
10GPT-5.247.17#105
11GPT-4o Mini47.08#588
12InternVL3-78B45.91#345
13Qwen 3 VL 235B A22B Instruct45.56#264
14Llama 3.2 90B Vision Instruct45.15#611
15Gemma 3 27B (IT)44.81#509

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=timespot-season · How It Works · Data refreshed daily, snapshot 2026-10-11.