ConflictQA - KG Conflict (Non-Complementary): leaderboard

Metric: Exact match (%) on 802 single-hop (non-complementary) questions when the provided evidence contains conflicting knowledge-graph triples (TripleConf); zero-shot with a conflict-aware prompt, the mean of two evidence orders; exact match of the generated answer set; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 12 models tracked.

Top models

#ModelScore
1O3 Mini38.03
2GPT-3.5 Turbo (0125)36.41
3DeepSeek V3.2 (Thinking)33.23
4GPT-4o32.67
5Llama 3.1 70B Instruct32.67
6GPT-5.132.05
7DeepSeek V3.2 (Non-reasoning)30.3
8Mistral Large 329.3
9Llama 3.1 8B Instruct19.2
10Qwen 3 8B7.61

Interactive version: theaggregate.ai/benchmark?slug=conflictqa-kg-conflict-non-complementary · How It Works · Data refreshed daily, snapshot 2026-10-07.