BizBench - TAT-QA Extract: leaderboard

Metric: Accuracy (%; 248 TAT-QA questions answerable by extracting a number from the text or table, 3-shot). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Llama 2 70B94.4
2Mixtral 8x7B91.1
3GPT-490.3
4Llama 2 13B88.7
5Mistral 7B87.9
6GPT-3.584.2
7Llama 2 7B83.1
8mpt-30B81.5
9falcon-40B80.6
10starcoder75
11falcon-7B62.9
12mpt-7B62.1

Interactive version: theaggregate.ai/benchmark?slug=bizbench-tat-qa-extract · How It Works · Data refreshed daily, snapshot 2026-09-26.