AnyGroundBench - Surgery TVG: leaderboard

Metric: tIoU@0.3 (%; share of queries whose temporal IoU is at least 0.3) on the 216 surgery-domain test queries (EgoSurgery and CholecTrack20 surgical videos), temporal video grounding (the queried event's start and end), zero-shot, default inference settings (Qwen3.5 with thinking disabled); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Flash37.9
2Gemini 3.1 Pro (Preview)32.8
3Gemini 2.5 Pro31.4
4Qwen 3.5 9B (Non-reasoning)28.2
5Qwen 3.5 4B (Non-reasoning)28.2
6Gemini 2.5 Flash22.6
7GPT-5.122.6
8GPT-4o13.4
9InternVL3-8B10.6
10InternVL3-14B10.6
11InternVL3.5-8B7.4

Interactive version: theaggregate.ai/benchmark?slug=anygroundbench-surgery-tvg · How It Works · Data refreshed daily, snapshot 2026-09-29.