CapRiCorn-1K - Referential Consistency: leaderboard

Metric: Subject referential consistency (%; share of same-subject description pairs a GPT-4.1 judge clusters as one referent, over all pairs of that subject's keypoints, audiovisual captions of 1,000 videos). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemini 3 Flash39.6
2Gemini 3.1 Pro (Preview)39.1
3Qwen3 Omni 30B A3B Instruct1.6
4Qwen2.5-Omni-7B0.6

Interactive version: theaggregate.ai/benchmark?slug=capricorn-1k-referential-consistency · How It Works · Data refreshed daily, snapshot 2026-09-26.