PharmaBatchDB NLQ-to-SQL: leaderboard

Metric: Factual consistency (%): Jaccard similarity between the result-row sets of the generated and the reference T-SQL query (1 when both are empty, 0 for a query that is not extracted or fails validation), times 100, averaged over 60 manufacturing questions over a synthetic MS SQL Server database (Batch, MES and CIP modules, 20 each; 18 easy, 24 medium, 18 hard), zero-shot T-SQL generation with the full module schema in the prompt, temperature 0.1, 4-bit quantized models served by Ollama on consumer CPU hardware; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1qwen2.5-coder:7b-instruct-q4_K_M (Ollama)34.08
2llama3.1:8b-instruct-q4_K_M (Ollama)30.38
3mistral:7b-instruct-q4_K_M (Ollama)13.33
4meditron:7b-q4_0 (Ollama)1.67

Interactive version: theaggregate.ai/benchmark?slug=pharmabatchdb-nlq-to-sql · How It Works · Data refreshed daily, snapshot 2026-09-29.