DailyClue - Daily Commonsense: leaderboard

Metric: Accuracy (%) on daily commonsense reasoning (food, health, customs, planning and consumption) on DailyClue, 666 question-image pairs from daily scenarios whose answer requires finding a decisive visual clue; question-clue-answer triplets drafted by GPT-5 and Gemini 2.5 Pro, manually verified, and kept only when at most two of o4-mini, Gemini 2.5 Flash and Claude 3.7 Sonnet answered correctly; exact match, with a Gemini 2.5 Pro judge for open-ended answers; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 24 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro62.77
2Gemini 2.5 Flash59.44
3O4 Mini58.33
4Qwen 3 VL 235B A22B (Thinking)56.67
5InternVL3-78B52.78
6GPT-551.67
7Qwen 3 VL 235B A22B Instruct50
8Claude Sonnet 4.549.44
9Claude Sonnet 448.89
10Qwen 2.5 VL 72B Instruct48.33
11InternVL3-38B47.22
12Claude 3.7 Sonnet47.22
13Qwen 2.5 VL 32B Instruct42.78
14Qwen 2.5 VL 7B Instruct37.22
15InternVL3-8B31.67

Interactive version: theaggregate.ai/benchmark?slug=dailyclue-daily-commonsense · How It Works · Data refreshed daily, snapshot 2026-10-07.