EgoSafetyBench: leaderboard

Metric: Video-level balanced accuracy (%; mean of the recall on unsafe videos and on safe videos over the 800 situational-track scenarios, a video called unsafe when any of its chunks is flagged, macro-averaged over videos; ten VLM safety guards judging egocentric robot videos split into ten half-second chunks (ten frames each at 1280x720), each chunk judged alone under the current-chunk protocol with a single-word verdict and greedy decoding; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Flash Lite87.3
2Qwen 3.5 4B85.5
3Gemini 3.5 Flash85.4
4Qwen 3 VL 4B Instruct84.9
5InternVL3.5-8B82.7
6Claude Sonnet 4.682.5
7Gemma 3 4B (IT)65.3
8Qwen 3.5 0.8B65.3

Interactive version: theaggregate.ai/benchmark?slug=egosafetybench · How It Works · Data refreshed daily, snapshot 2026-09-29.