BIRD-History: leaderboard

Metric: Execution accuracy (%; 1,393 underspecified text-to-SQL tasks over 11 databases whose implicit business knowledge appears only in historical SQL logs; foundation model generating SQL directly without a history retriever, temperature 0). Source: arxiv.org. Saturation forecast: Around August 2028. 5 models tracked.

Top models

#ModelScore
1Gemini 3 Flash33.24
2GPT-5.3 Codex28.93
3DeepSeek V3.221.11
4Llama 3.3 70B Instruct16.65

Interactive version: theaggregate.ai/benchmark?slug=bird-history · How It Works · Data refreshed daily, snapshot 2026-09-29.