PharmaBatchDB NLQ-to-SQL: leaderboard
Metric: Factual consistency (%): Jaccard similarity between the result-row sets of the generated and the reference T-SQL query (1 when both are empty, 0 for a query that is not extracted or fails validation), times 100, averaged over 60 manufacturing questions over a synthetic MS SQL Server database (Batch, MES and CIP modules, 20 each; 18 easy, 24 medium, 18 hard), zero-shot T-SQL generation with the full module schema in the prompt, temperature 0.1, 4-bit quantized models served by Ollama on consumer CPU hardware; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | qwen2.5-coder:7b-instruct-q4_K_M (Ollama) | 34.08 |
| 2 | llama3.1:8b-instruct-q4_K_M (Ollama) | 30.38 |
| 3 | mistral:7b-instruct-q4_K_M (Ollama) | 13.33 |
| 4 | meditron:7b-q4_0 (Ollama) | 1.67 |
Interactive version: theaggregate.ai/benchmark?slug=pharmabatchdb-nlq-to-sql · How It Works · Data refreshed daily, snapshot 2026-09-29.