ForesightSafety Bench (Embodied AI Safety) - Collaborative Norm Safety: leaderboard

Metric: Violation rate (%) over the collaborative norm safety scenarios: unsafe robot task instructions from the paper's ForesightSafetyBench-EmbodiedAI-O set, given to the model as a robot task planner; a violation is an unsafe instruction the model executes instead of refusing, judged by gpt-4o-2024-11-20 (the paper's primary evaluator) with a PAIR-style rating prompt; lower is better. Source: arxiv.org. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash (Preview)30.86#78
2Kimi K2.533.33#139
3Gemini 3 Pro (Preview)33.33#64
4GPT-5.240.74#105
5Gemini 2.5 Flash44.44#237
6Claude 3.5 Haiku49.38#553
7Grok 4 Fast (Non-reasoning)49.38#242 (Grok 4 Fast)
8GPT-4o (2024-11-20)55.56#369
9GLM-4.756.79#185
10Seed-1.658.02#257
11Qwen 3 235B A22B 2507 Instruct59.26#291
12Grok 460.49#169
13Qwen 2.5 72B Instruct61.73#436
14Kimi K2 090561.73#282
15Seed 1.861.73#136

Interactive version: theaggregate.ai/benchmark?slug=foresightsafety-bench-embodied-ai-safety-collaborative-norm-safety · How It Works · Data refreshed daily, snapshot 2026-10-11.