V-DyKnow (Visual Prompt): leaderboard
Metric: Correct rate (%): the share of responses giving the currently valid attribute (rather than an outdated or irrelevant one) on the 139 time-sensitive V-DyKnow facts (82 about countries, 28 about athletes, 29 about organizations; attributes checked against Wikidata as of November 2025), each asked with three prompt lexicalizations and scored by the best of the three (upper bound), greedy decoding, answer with the name only, with the entity shown as an image (a flag and coat of arms, a portrait or a logo) in place of its name; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.1 (2025-11-13) | 75 | #97 |
| 2 | GPT-4.1 | 71 | #240 |
| 3 | Qwen 2.5 VL 7B Instruct | 32 | #643 |
| 4 | Qwen 2 VL 7B Instruct | 28 | #816 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=v-dyknow-visual-prompt · How It Works · Data refreshed daily, snapshot 2026-10-11.