VideoFDB - Perception: leaderboard
Metric: gpt-4o judge score (0-5; the 105 held-out test clips of the seven nonverbal conversational dynamics scored for perception (237 clips of 11 dynamics in the full benchmark), streamed in real time; the judge reads the transcribed, time-stamped agent response with the clip annotations; mean of the fluency, conversational-flow and semantic-grounding rubrics; user audio and video streamed, video at 1 FPS (Mini-Omni2: one frame per user speech segment)). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | MiniCPM-o-4.5 | 3.4 |
Interactive version: theaggregate.ai/benchmark?slug=videofdb-perception · How It Works · Data refreshed daily, snapshot 2026-09-26.