AnyGroundBench - Public Security TVG: leaderboard
Metric: tIoU@0.3 (%; share of queries whose temporal IoU is at least 0.3) on the 250 public security-domain test queries (UCA surveillance and DoTA traffic-accident videos), temporal video grounding (the queried event's start and end), zero-shot, default inference settings (Qwen3.5 with thinking disabled); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 69.4 |
| 2 | Gemini 3 Flash | 66.8 |
| 3 | Gemini 2.5 Pro | 65.8 |
| 4 | GPT-5.1 | 61.2 |
| 5 | GPT-4o | 56 |
| 6 | Gemini 2.5 Flash | 51.2 |
| 7 | Qwen 3.5 9B (Non-reasoning) | 50.4 |
| 8 | Qwen 3.5 4B (Non-reasoning) | 49.2 |
| 9 | InternVL3-14B | 13.2 |
| 10 | InternVL3-8B | 4.4 |
| 11 | InternVL3.5-8B | 3.6 |
Interactive version: theaggregate.ai/benchmark?slug=anygroundbench-public-security-tvg · How It Works · Data refreshed daily, snapshot 2026-09-29.