BizBench - CodeFinQA: leaderboard

Metric: Accuracy (%; 844 FinQA questions over report text and tables answered by generating Python code, answer within 1% of the reference, 3-shot). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-478.8
2GPT-3.567.5
3Mixtral 8x7B58.5
4Llama 2 70B57.3
5Mistral 7B48.8
6Llama 2 13B33.4
7starcoder31.2
8mpt-30B31
9Llama 2 7B21.9
10falcon-40B18.4
11mpt-7B6.6
12falcon-7B2

Interactive version: theaggregate.ai/benchmark?slug=bizbench-codefinqa · How It Works · Data refreshed daily, snapshot 2026-09-26.