V-DyKnow (Textual Prompt): leaderboard

Metric: Correct rate (%): the share of responses giving the currently valid attribute (rather than an outdated or irrelevant one) on the 139 time-sensitive V-DyKnow facts (82 about countries, 28 about athletes, 29 about organizations; attributes checked against Wikidata as of November 2025), each asked with three prompt lexicalizations and scored by the best of the three (upper bound), greedy decoding, answer with the name only, with the entity named in the text of the question; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.1 (2025-11-13)76#97
2GPT-4.172#240
3Qwen 2.5 VL 7B Instruct39#643
4Qwen 2 VL 7B Instruct36#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=v-dyknow-textual-prompt · How It Works · Data refreshed daily, snapshot 2026-10-11.