ConflictBench (Multi-Modal): leaderboard
Metric: Task success rate (%): share of scenarios in which the agent's actions reach a human-favorable terminal outcome in the environment, over ConflictBench's 150 multi-turn human-AI conflict scenarios (51 self-preservation versus human safety, 47 resource prioritization and 52 deceptive alignment episodes, adapted from PacifAIst), macro-averaged over the three categories, temperature 0; multi-modal setting: the agent also receives generated video of the world state from a visual world model (frames at 1 fps), at most 10 turns; higher is better. Source: arxiv.org. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 | 70.05 | #91 |
| 2 | GPT-4o | 60.93 | #333 |
| 3 | Qwen 3 VL 30B A3B Instruct | 46.7 | #365 |
| 4 | Gemini 2.5 Flash | 43.68 | #237 |
Interactive version: theaggregate.ai/benchmark?slug=conflictbench-multi-modal · How It Works · Data refreshed daily, snapshot 2026-10-11.