EgoSafetyBench - Sequential Detection: leaderboard

Metric: Caught rate (%; share of unsafe situational videos whose first alarm falls within one half-second chunk after the annotated hazard onset, chunks streamed in order; ten VLM safety guards judging egocentric robot videos split into ten half-second chunks (ten frames each at 1280x720), each chunk judged alone under the current-chunk protocol with a single-word verdict and greedy decoding; higher is better). Source: arxiv.org. Saturation forecast: Around March 2028. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.657.4
2Gemini 3.5 Flash54.2
3Gemini 3.1 Flash Lite52.3
4Qwen 3 VL 4B Instruct47.7
5Qwen 3.5 4B45.6
6Gemma 3 4B (IT)43.5
7InternVL3.5-8B39.2
8Qwen 3.5 0.8B15.8

Interactive version: theaggregate.ai/benchmark?slug=egosafetybench-sequential-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.