ConflictQA - KG Conflict (Complementary): leaderboard

Metric: Exact match (%) on 430 complementary (answers need both sources) questions when the provided evidence contains conflicting knowledge-graph triples (TripleConf); zero-shot with a conflict-aware prompt, the mean of two evidence orders; exact match of the generated answer set; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 12 models tracked.

Top models

#ModelScore
1O3 Mini37.67
2DeepSeek V3.2 (Thinking)36.98
3GPT-4o33.72
4GPT-5.133.72
5Llama 3.1 70B Instruct31.16
6GPT-3.5 Turbo (0125)30.7
7Mistral Large 329.77
8DeepSeek V3.2 (Non-reasoning)26.74
9Llama 3.1 8B Instruct17.67
10Qwen 3 8B8.14

Interactive version: theaggregate.ai/benchmark?slug=conflictqa-kg-conflict-complementary · How It Works · Data refreshed daily, snapshot 2026-10-07.