HomeSafe-Bench: leaderboard

Metric: Weighted Safety Score (0-100) on the 438 hazardous household videos (Veo-3.1 generations and BEHAVIOR simulations, 10 FPS with an overlaid timestamp, read in 2-second sliding windows): each video scores 100 when the model's warning falls in the optimal window (intent onset to the intervention deadline 200 ms before the point of no return), less for later warnings, and 0 for premature or missed warnings, averaged over videos; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 16 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 8B Instruct24.77#401
2GPT-5.121.12#131
3Claude Opus 4.118.38#132
4Qwen 3 VL 4B Instruct16.27#506

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=homesafe-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.