ConflictQA - Text Conflict (Non-Complementary): leaderboard

Metric: Exact match (%) on 802 single-hop (non-complementary) questions when the provided evidence contains conflicting text passages (TextConf); zero-shot with a conflict-aware prompt, the mean of two evidence orders; exact match of the generated answer set; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 12 models tracked.

Top models

#ModelScore
1DeepSeek V3.2 (Non-reasoning)53.43
2O3 Mini52.56
3Qwen 3 8B52.44
4GPT-5.151.88
5Llama 3.1 70B Instruct48.96
6GPT-4o46.57
7GPT-3.5 Turbo (0125)43.77
8Llama 3.1 8B Instruct41.15
9DeepSeek V3.2 (Thinking)40.15
10Mistral Large 332.17

Interactive version: theaggregate.ai/benchmark?slug=conflictqa-text-conflict-non-complementary · How It Works · Data refreshed daily, snapshot 2026-10-07.