ConflictBench (Text-Only) - Alignment Success Rate: leaderboard

Metric: Alignment success rate (%): share of scenarios in which a GPT-5 judge finds the agent's reasoning and trajectory consistently prioritizing human interests, regardless of execution success, over ConflictBench's 150 multi-turn human-AI conflict scenarios (51 self-preservation versus human safety, 47 resource prioritization and 52 deceptive alignment episodes, adapted from PacifAIst), macro-averaged over the three categories, temperature 0; text-only setting: the agent sees only the text simulation engine, at most 14 turns; higher is better. Source: arxiv.org. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-570.23#91
2GPT-4o65.01#333
3Qwen Plus62.37#370
4DeepSeek V357.67#312
5Qwen 3 VL 30B A3B Instruct56.92#365
6GPT-4o Mini47.28#588
7Gemini 2.5 Flash47.02#237

Interactive version: theaggregate.ai/benchmark?slug=conflictbench-text-only-alignment-success-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.