BI-Bench (SQL): leaderboard

Metric: Accuracy (%; 100 end-to-end business-intelligence queries curated from real Power BI dashboards (82 projects, raw source tables needing search, reshaping, joins and analysis); an answer is correct when the output table matches the ground-truth table in shape and cell values up to row and column order, with numeric tolerance; mean over 10 runs; SQL setting: tables loaded into SQLite and the model iterates with a generic code-execution tool only). Source: arxiv.org. Saturation forecast: Around October 2027. 10 models tracked.

Top models

#ModelScore
1O4 Mini48.2
2GPT-5.546.7
3GPT-5.246.1
4Mistral Large 336.2
5Llama 4 Maverick27.5
6GPT-4o27.1
7DeepSeek V4 Pro23.5
8Kimi K2.620
9GPT-OSS-120B5.3
10Qwen 3 8B5.2

Interactive version: theaggregate.ai/benchmark?slug=bi-bench-sql · How It Works · Data refreshed daily, snapshot 2026-09-26.