Text2GraphQuery-Bench (Zero-Shot) - GQL: leaderboard

Metric: Execution accuracy (%) of the generated ISO GQL query with zero-shot prompting; Text2GraphQuery-Bench's 2,783 test questions executable on all three engines (6 held-out graph databases in 5 domains), the model given the graph schema and the question, every model in non-thinking mode; the predicted query is executed (Cypher on TuGraph-DB, GQL on Spanner Graph, SQL/PGQ on Oracle Database) and counts when its result equals the gold query's; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3.7 Max (Non-reasoning)46.9#71 (Qwen 3.7 Max)
2GPT-5.5 (Non-reasoning)46.2#26 (GPT-5.5)
3Claude Opus 4.845.8#33
4Kimi K2.6 (Non-reasoning)45.3#99 (Kimi K2.6)
5DeepSeek V4 Pro (Non-reasoning)40.3#96 (DeepSeek V4 Pro)
6Qwen 3 8B (Non-reasoning)16#667 (Qwen 3 8B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=text2graphquery-bench-zero-shot-gql · How It Works · Data refreshed daily, snapshot 2026-10-11.