MedVidBench — leaderboard
Medical and surgical video understanding benchmark for video large language models, covering 6,245 test samples across eight tasks including temporal action localization, spatiotemporal grounding, captioning, next-action prediction, CVS assessment, video summary, region captioning, and surgical skill assessment.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=medvidbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.