ConflictBench (Text-Only): leaderboard

Metric: Task success rate (%): share of scenarios in which the agent's actions reach a human-favorable terminal outcome in the environment, over ConflictBench's 150 multi-turn human-AI conflict scenarios (51 self-preservation versus human safety, 47 resource prioritization and 52 deceptive alignment episodes, adapted from PacifAIst), macro-averaged over the three categories, temperature 0; text-only setting: the agent sees only the text simulation engine, at most 14 turns; higher is better. Source: arxiv.org. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-565.51#91
2Qwen Plus59.71#370
3GPT-4o59.57#333
4DeepSeek V353.64#312
5Qwen 3 VL 30B A3B Instruct53.57#365
6Gemini 2.5 Flash44.35#237
7GPT-4o Mini35.1#588

Interactive version: theaggregate.ai/benchmark?slug=conflictbench-text-only · How It Works · Data refreshed daily, snapshot 2026-10-11.