BIRD-Python: leaderboard
Metric: LLM-validated execution accuracy (%; generated pandas code executed against the corrected ('verified') 1,534 BIRD dev questions (925 simple, 464 moderate, 145 challenging), an answer counting when an LLM validator judges its output equivalent to the reference result; one generation per question). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 Max | 63.43 |
| 2 | DeepSeek R1 | 62.52 |
| 3 | Qwen 3 Coder Plus | 61.6 |
| 4 | Qwen Max | 60.82 |
| 5 | Qwen 3 32B (Thinking) | 60.1 |
| 6 | Qwen 3 14B (Reasoning) | 47.98 |
| 7 | Qwen 3 Coder 30B A3B Instruct | 45.83 |
Interactive version: theaggregate.ai/benchmark?slug=bird-python · How It Works · Data refreshed daily, snapshot 2026-09-29.