AnyGroundBench - Surgery TVG: leaderboard
Metric: tIoU@0.3 (%; share of queries whose temporal IoU is at least 0.3) on the 216 surgery-domain test queries (EgoSurgery and CholecTrack20 surgical videos), temporal video grounding (the queried event's start and end), zero-shot, default inference settings (Qwen3.5 with thinking disabled); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 37.9 |
| 2 | Gemini 3.1 Pro (Preview) | 32.8 |
| 3 | Gemini 2.5 Pro | 31.4 |
| 4 | Qwen 3.5 9B (Non-reasoning) | 28.2 |
| 5 | Qwen 3.5 4B (Non-reasoning) | 28.2 |
| 6 | Gemini 2.5 Flash | 22.6 |
| 7 | GPT-5.1 | 22.6 |
| 8 | GPT-4o | 13.4 |
| 9 | InternVL3-8B | 10.6 |
| 10 | InternVL3-14B | 10.6 |
| 11 | InternVL3.5-8B | 7.4 |
Interactive version: theaggregate.ai/benchmark?slug=anygroundbench-surgery-tvg · How It Works · Data refreshed daily, snapshot 2026-09-29.