DuckDB-NSQL — leaderboard

SQL generation benchmark using DuckDB: 75 questions across easy, medium, and hard difficulty testing natural-language-to-SQL translation with execution accuracy.

Metric: Execution Accuracy (%). Source: huggingface.co. Status: saturation imminent. 99 models tracked.

Top models

#ModelScore
1DeepSeek V3 Chat80
2Gemini 2.5 Flash (Preview 05-20)77.3
3GPT-4.174.7
4Claude Sonnet 4.574.7
5DeepSeek R174.7
6GPT-4o (2024-11-20)74.7
7Claude 3.7 Sonnet73.3
8O3 Mini72
9GPT-OSS-120B70.7
10Mistral Small 370.7
11O1 Preview70.7
12DeepSeek V3 (0324)69.3
13O169.3
14Claude 3.5 Sonnet68
15Qwen 2.5 Coder 32B Instruct68

Interactive version: theaggregate.ai/benchmark?slug=duckdb-nsql · How the rankings work · Data refreshed daily, snapshot 2026-07-22.