OmniStarPro-RNG (Offline): leaderboard

Metric: Semantic correctness (0-10; GPT-4o judge score of each narration against the ground-truth clip caption, mean of semantic accuracy, language quality and information completeness; real-time narration generation with responses decoded at prescribed timestamps, 1,000 evaluation streams of the OmniStarPro-Live partition). Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScore
1GPT-4o5.03
2MiniCPM-V-2.64.34

Interactive version: theaggregate.ai/benchmark?slug=omnistarpro-rng-offline · How It Works · Data refreshed daily, snapshot 2026-09-26.