TKFQA - Reasoning-Chain Accuracy (Table-Text-KG Order): leaderboard

Metric: Reasoning-chain accuracy (%; entity-level agreement of the generated reasoning chain with the annotated counterfactual chain, for counterfactual multi-hop questions grounded in a table, text passages and a knowledge graph, 1,012 test items; contexts given in table, text, knowledge-graph order; mean of three runs). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1O379.05
2Gemini 2.5 Flash77.93
3DeepSeek V3.269.2
4O4 Mini68.18
5DeepSeek V4 Flash65.26
6Kimi K2.564.04
7MiniMax-M2.755.34
8GPT-551.52
9Qwen 3 30B A3B43.31
10Gemini 2.5 Flash Lite37.25
11Grok 4.332.61
12GPT-4.1 Mini29.59
13Qwen 3 8B20.14
14Gemma 3 12B19.95

Interactive version: theaggregate.ai/benchmark?slug=tkfqa-reasoning-chain-accuracy-table-text-kg-order · How It Works · Data refreshed daily, snapshot 2026-09-26.