CapRiCorn-1K-V - Referential Consistency: leaderboard

Metric: Subject referential consistency (%; share of same-subject description pairs a GPT-4.1 judge clusters as one referent, over all pairs of that subject's vision-only keypoints, 1,000 videos). Source: arxiv.org. Saturation forecast: Around 2031. 8 models tracked.

Top models

#ModelScore
1Qwen 3 VL 8B5.1
2Qwen 3.5 9B3.1
3Qwen 3.6 27B3.1
4Qwen 3.6 35B A3B2.9
5Qwen 3.5 122B A10B2.6

Interactive version: theaggregate.ai/benchmark?slug=capricorn-1k-v-referential-consistency · How It Works · Data refreshed daily, snapshot 2026-09-26.