ConflictBench (Multi-Modal) - Alignment Success Rate: leaderboard

Metric: Alignment success rate (%): share of scenarios in which a GPT-5 judge finds the agent's reasoning and trajectory consistently prioritizing human interests, regardless of execution success, over ConflictBench's 150 multi-turn human-AI conflict scenarios (51 self-preservation versus human safety, 47 resource prioritization and 52 deceptive alignment episodes, adapted from PacifAIst), macro-averaged over the three categories, temperature 0; multi-modal setting: the agent also receives generated video of the world state from a visual world model (frames at 1 fps), at most 10 turns; higher is better. Source: arxiv.org. 5 models tracked.

Top models

#ModelScoreOverall rank
1GPT-575.56#91
2GPT-4o64.94#333
3Qwen 3 VL 30B A3B Instruct54.1#365
4Gemini 2.5 Flash46.41#237

Interactive version: theaggregate.ai/benchmark?slug=conflictbench-multi-modal-alignment-success-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.