BizBench - ConvFinQA Extract: leaderboard

Metric: Accuracy (%; 916 ConvFinQA questions answerable by extracting a number from the report text or table, 3-shot). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-494
2Llama 2 70B92.8
3Mixtral 8x7B92.8
4GPT-3.592.4
5Mistral 7B91.5
6Llama 2 13B88.2
7Llama 2 7B86.1
8mpt-30B85.5
9falcon-40B82.8
10starcoder79.7
11mpt-7B71.7
12falcon-7B66.5

Interactive version: theaggregate.ai/benchmark?slug=bizbench-convfinqa-extract · How It Works · Data refreshed daily, snapshot 2026-09-26.