BIRD-CRITIC — leaderboard

SQL debugging and problem-solving benchmark evaluating LLMs on diagnosing and resolving real-world user SQL issues.

Metric: Score. Source: github.com. Status: saturation imminent. 6 models tracked.

Top models

#ModelScore
1O3 Mini (2025-01-31)34.5
2DeepSeek Reasoner33.67
3O1 Preview (2024-09-12)33.33
4Claude 3.7 Sonnet (20250219) (Thinking)30.67
5Gemini 2.0 Flash (01-21) (Thinking)30.17
6Grok 329.83

Interactive version: theaggregate.ai/benchmark?slug=bird-critic · How the rankings work · Data refreshed daily, snapshot 2026-07-22.