BizBench - CodeTAT-QA: leaderboard

Metric: Accuracy (%; 392 TAT-QA table questions answered by generating Python over a dataframe, answer within 1% of the reference, 3-shot). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-490.6
2GPT-3.587.6
3Mixtral 8x7B83.9
4Llama 2 70B79.1
5Mistral 7B75
6starcoder70.2
7Llama 2 13B65.1
8mpt-30B64.8
9falcon-40B38.5
10Llama 2 7B37
11mpt-7B30.4
12falcon-7B7.4

Interactive version: theaggregate.ai/benchmark?slug=bizbench-codetat-qa · How It Works · Data refreshed daily, snapshot 2026-09-26.