Text2GraphQuery-Bench (Few-Shot) - GQL: leaderboard
Metric: Execution accuracy (%) of the generated ISO GQL query with three fixed training-set demonstrations prepended to every question; Text2GraphQuery-Bench's 2,783 test questions executable on all three engines (6 held-out graph databases in 5 domains), the model given the graph schema and the question, every model in non-thinking mode; the predicted query is executed (Cypher on TuGraph-DB, GQL on Spanner Graph, SQL/PGQ on Oracle Database) and counts when its result equals the gold query's; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Opus 4.8 | 58.2 | #33 |
| 2 | GPT-5.5 (Non-reasoning) | 54.9 | #26 (GPT-5.5) |
| 3 | Qwen 3.7 Max (Non-reasoning) | 54.8 | #71 (Qwen 3.7 Max) |
| 4 | Kimi K2.6 (Non-reasoning) | 51.1 | #99 (Kimi K2.6) |
| 5 | DeepSeek V4 Pro (Non-reasoning) | 45.4 | #96 (DeepSeek V4 Pro) |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=text2graphquery-bench-few-shot-gql · How It Works · Data refreshed daily, snapshot 2026-10-11.