Text2GraphQuery-Bench (Zero-Shot) - GQL: leaderboard
Metric: Execution accuracy (%) of the generated ISO GQL query with zero-shot prompting; Text2GraphQuery-Bench's 2,783 test questions executable on all three engines (6 held-out graph databases in 5 domains), the model given the graph schema and the question, every model in non-thinking mode; the predicted query is executed (Cypher on TuGraph-DB, GQL on Spanner Graph, SQL/PGQ on Oracle Database) and counts when its result equals the gold query's; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Qwen 3.7 Max (Non-reasoning) | 46.9 | #71 (Qwen 3.7 Max) |
| 2 | GPT-5.5 (Non-reasoning) | 46.2 | #26 (GPT-5.5) |
| 3 | Claude Opus 4.8 | 45.8 | #33 |
| 4 | Kimi K2.6 (Non-reasoning) | 45.3 | #99 (Kimi K2.6) |
| 5 | DeepSeek V4 Pro (Non-reasoning) | 40.3 | #96 (DeepSeek V4 Pro) |
| 6 | Qwen 3 8B (Non-reasoning) | 16 | #667 (Qwen 3 8B) |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=text2graphquery-bench-zero-shot-gql · How It Works · Data refreshed daily, snapshot 2026-10-11.