BI-Bench (Python): leaderboard

Metric: Accuracy (%; 100 end-to-end business-intelligence queries curated from real Power BI dashboards (82 projects, raw source tables needing search, reshaping, joins and analysis); an answer is correct when the output table matches the ground-truth table in shape and cell values up to row and column order, with numeric tolerance; mean over 10 runs; Python setting: tables given as CSV files and the model iterates with a generic code-execution tool only). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1GPT-5.559.5
2O4 Mini57.4
3Kimi K2.655.6
4DeepSeek V4 Pro53.2
5GPT-5.248.6
6Mistral Large 335
7GPT-4o30.8
8Llama 4 Maverick28.6
9Qwen 3 8B5.4
10GPT-OSS-120B3.1

Interactive version: theaggregate.ai/benchmark?slug=bi-bench-python · How It Works · Data refreshed daily, snapshot 2026-09-26.