DailyClue - Location Identification: leaderboard

Metric: Accuracy (%) on location identification (country and first-level region of the scene) on DailyClue, 666 question-image pairs from daily scenarios whose answer requires finding a decisive visual clue; question-clue-answer triplets drafted by GPT-5 and Gemini 2.5 Pro, manually verified, and kept only when at most two of o4-mini, Gemini 2.5 Flash and Claude 3.7 Sonnet answered correctly; exact match, with a Gemini 2.5 Pro judge for open-ended answers; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 24 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro41.5
2GPT-538
3Gemini 2.5 Flash32.5
4O4 Mini25.5
5Qwen 2.5 VL 72B Instruct24.5
6Qwen 3 VL 235B A22B (Thinking)23
7Qwen 3 VL 235B A22B Instruct22.5
8Claude Sonnet 422
9Qwen 2.5 VL 32B Instruct21.5
10Claude Sonnet 4.521
11Claude 3.7 Sonnet18.5
12InternVL3-78B18
13InternVL3-38B17
14Qwen 2.5 VL 7B Instruct15
15InternVL3-8B13.5

Interactive version: theaggregate.ai/benchmark?slug=dailyclue-location-identification · How It Works · Data refreshed daily, snapshot 2026-10-07.