HomeSafe-Bench: leaderboard
Metric: Weighted Safety Score (0-100) on the 438 hazardous household videos (Veo-3.1 generations and BEHAVIOR simulations, 10 FPS with an overlaid timestamp, read in 2-second sliding windows): each video scores 100 when the model's warning falls in the optimal window (intent onset to the intervention deadline 200 ms before the point of no return), less for later warnings, and 0 for premature or missed warnings, averaged over videos; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 16 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Qwen 3 VL 8B Instruct | 24.77 | #401 |
| 2 | GPT-5.1 | 21.12 | #131 |
| 3 | Claude Opus 4.1 | 18.38 | #132 |
| 4 | Qwen 3 VL 4B Instruct | 16.27 | #506 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=homesafe-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.