TKFQA (Table-Text-KG Order): leaderboard

Metric: Exact match (%; final answers to counterfactual multi-hop questions grounded in a table, text passages and a knowledge graph, 1,012 test items; contexts given in table, text, knowledge-graph order; mean of three runs). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1DeepSeek V4 Flash84.66
2O384.5
3Gemini 2.5 Flash84.2
4DeepSeek V3.283.05
5O4 Mini82.5
6Kimi K2.579.31
7Grok 4.379.22
8GPT-575.38
9MiniMax-M2.772.42
10GPT-4.1 Mini71.44
11Gemini 2.5 Flash Lite69.93
12Qwen 3 30B A3B64.9
13Gemma 3 12B62.33
14Qwen 3 8B52.35

Interactive version: theaggregate.ai/benchmark?slug=tkfqa-table-text-kg-order · How It Works · Data refreshed daily, snapshot 2026-09-26.