ForesightSafety Bench (Embodied AI Safety) - Living-Being Physical Contact Safety: leaderboard

Metric: Violation rate (%) over the living-being physical contact safety scenarios: unsafe robot task instructions from the paper's ForesightSafetyBench-EmbodiedAI-O set, given to the model as a robot task planner; a violation is an unsafe instruction the model executes instead of refusing, judged by gpt-4o-2024-11-20 (the paper's primary evaluator) with a PAIR-style rating prompt; lower is better. Source: arxiv.org. 22 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.211.11#105
2Gemini 2.5 Flash15.28#237
3Gemini 3 Flash (Preview)16.11#78
4GLM-4.741.94#185
5Grok 4 Fast (Non-reasoning)52.92#242 (Grok 4 Fast)
6Qwen 3 235B A22B 2507 Instruct53.89#291
7Kimi K2 090554.58#282
8Llama 4 Maverick55.69#451
9GPT-4o (2024-11-20)58.19#369
10Llama 3.3 70B Instruct59.03#520
11Claude 3.5 Haiku59.31#553
12Grok 459.58#169
13Qwen 2.5 72B Instruct60#436
14Kimi K2.561.67#139
15DeepSeek V3.261.67#198

Interactive version: theaggregate.ai/benchmark?slug=foresightsafety-bench-embodied-ai-safety-living-being-physical-contact-safety · How It Works · Data refreshed daily, snapshot 2026-10-11.