Text2GraphQuery-Bench (Few-Shot) - SQL/PGQ: leaderboard

Metric: Execution accuracy (%) of the generated SQL/PGQ query with three fixed training-set demonstrations prepended to every question; Text2GraphQuery-Bench's 2,783 test questions executable on all three engines (6 held-out graph databases in 5 domains), the model given the graph schema and the question, every model in non-thinking mode; the predicted query is executed (Cypher on TuGraph-DB, GQL on Spanner Graph, SQL/PGQ on Oracle Database) and counts when its result equals the gold query's; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.5 (Non-reasoning)51.2#26 (GPT-5.5)
2Claude Opus 4.850.9#33
3Kimi K2.6 (Non-reasoning)27.2#99 (Kimi K2.6)
4Qwen 3.7 Max (Non-reasoning)20.9#71 (Qwen 3.7 Max)
5DeepSeek V4 Pro (Non-reasoning)16.2#96 (DeepSeek V4 Pro)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=text2graphquery-bench-few-shot-sql-pgq · How It Works · Data refreshed daily, snapshot 2026-10-11.