SciTaRC: leaderboard

Metric: LLM-judge accuracy (%) with chain-of-thought prompting on the 371 expert-authored SciTaRC questions over tables from recent arXiv papers (LaTeX table sources as input), zero-shot; an answer counts only when a Llama-3.3-70B-Instruct judge gives full credit against the gold answer (partial matches count as wrong); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 21 models tracked.

Top models

#ModelScoreOverall rank
1GPT-576.8#91
2Grok 4.1 Fast (Reasoning)76.8#208 (Grok 4.1 Fast)
3Grok 474.4#169
4DeepSeek V3.2 (Thinking)73.6#198 (DeepSeek V3.2)
5Kimi K2 (Thinking)73.3#236 (Kimi K2)
6Claude Opus 4.572#79
7DeepSeek V3.2 (Non-reasoning)69.3#198 (DeepSeek V3.2)
8GPT-OSS-120B62.5#330
9Gemini 2.5 Flash60.6#237
10Qwen 3 Next 80B A3B49.3#306
11Kimi K249.1#236
12GPT-4o46.9#333
13Qwen 2.5 72B Instruct42#436
14DeepSeek R1 Distill Qwen 32B41.5#640
15Llama 3.3 70B Instruct34.5#520

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=scitarc · How It Works · Data refreshed daily, snapshot 2026-10-11.