ForesightSafety Bench (Embodied AI Safety) - Path Planning Safety: leaderboard

Metric: Violation rate (%) over the path planning safety scenarios: unsafe robot task instructions from the paper's ForesightSafetyBench-EmbodiedAI-O set, given to the model as a robot task planner; a violation is an unsafe instruction the model executes instead of refusing, judged by gpt-4o-2024-11-20 (the paper's primary evaluator) with a PAIR-style rating prompt; lower is better. Source: arxiv.org. 22 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.23.43#105
2Gemini 2.5 Flash7.3#237
3Gemini 3 Flash (Preview)10.3#78
4GLM-4.741.2#185
5Qwen 3 235B A22B 2507 Instruct49.79#291
6Llama 3.3 70B Instruct54.51#520
7Llama 4 Maverick54.94#451
8Kimi K2 090556.22#282
9Qwen 2.5 72B Instruct57.08#436
10GPT-4o (2024-11-20)60.09#369
11DeepSeek V3.260.52#198
12Grok 4 Fast (Non-reasoning)60.52#242 (Grok 4 Fast)
13Kimi K2.561.8#139
14Grok 462.66#169
15Seed-1.664.38#257

Interactive version: theaggregate.ai/benchmark?slug=foresightsafety-bench-embodied-ai-safety-path-planning-safety · How It Works · Data refreshed daily, snapshot 2026-10-11.