ConflictQA - Text Conflict (Complementary): leaderboard

Metric: Exact match (%) on 430 complementary (answers need both sources) questions when the provided evidence contains conflicting text passages (TextConf); zero-shot with a conflict-aware prompt, the mean of two evidence orders; exact match of the generated answer set; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 12 models tracked.

Top models

#ModelScore
1GPT-5.155.58
2DeepSeek V3.2 (Non-reasoning)51.86
3O3 Mini48.6
4Qwen 3 8B46.98
5GPT-4o44.42
6GPT-3.5 Turbo (0125)41.4
7DeepSeek V3.2 (Thinking)39.3
8Llama 3.1 70B Instruct33.72
9Mistral Large 328.84
10Llama 3.1 8B Instruct21.63

Interactive version: theaggregate.ai/benchmark?slug=conflictqa-text-conflict-complementary · How It Works · Data refreshed daily, snapshot 2026-10-07.