BIRD-History: leaderboard
Metric: Execution accuracy (%; 1,393 underspecified text-to-SQL tasks over 11 databases whose implicit business knowledge appears only in historical SQL logs; foundation model generating SQL directly without a history retriever, temperature 0). Source: arxiv.org. Saturation forecast: Around August 2028. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 33.24 |
| 2 | GPT-5.3 Codex | 28.93 |
| 3 | DeepSeek V3.2 | 21.11 |
| 4 | Llama 3.3 70B Instruct | 16.65 |
Interactive version: theaggregate.ai/benchmark?slug=bird-history · How It Works · Data refreshed daily, snapshot 2026-09-29.