ProLLM - SQL Disambiguation — leaderboard

ProLLM benchmark measuring LLM ability to disambiguate and generate correct SQL from ambiguous natural language queries.

Metric: Score (%). Source: www.prollm.ai. Status: saturation imminent. 52 models tracked.

Top models

#ModelScore
1GPT-4.553.1
2O149.2
3Grok 247.9
4GPT-3.5 Turbo47.7
5GPT-4.146.4
6Qwen 2.5 72B Instruct46
7DeepSeek V344.7
8Gemini 2.0 Flash42
9Mistral Small 3.140.8
10Command A40.6
11Grok 340.3
12GPT-4.1 Mini40.1
13GPT-4 Turbo37.5
14Nova Pro37.4
15GPT-4o36.7

Interactive version: theaggregate.ai/benchmark?slug=prollm-sql-disambiguation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.