UniQL - T-SQL: leaderboard

Metric: Execution accuracy (%) on the T-SQL dialect of UniQL (1,534 BIRD development-set questions with human-verified SQL in 16 aligned dialects): given the target dialect, schema and question, the model writes one query whose execution result must match the gold result; inference-only, one run; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.556.26
2Gemini 2.5 Pro55.61
3GPT-5.1 Codex54.63
4GPT-5 Mini53.85
5Qwen 3 32B51.69
6DeepSeek V4 Flash51.43
7Qwen 3 8B48.57
8Qwen 3 4B46.41
9Llama 3 70B Instruct40.61
10GPT-3.5 Turbo37.74
11Qwen 3 1.7B32.66
12Llama 3 8B Instruct21.12

Interactive version: theaggregate.ai/benchmark?slug=uniql-t-sql · How It Works · Data refreshed daily, snapshot 2026-09-29.