DataClawEval - PrestoSQL-Trino: leaderboard
Metric: Rule-based task score (%; per task a weighted sum, usually 0.7 and 0.3, of the artifact score, checked row by row against live databases by task-specific deterministic scripts, and the process score for exploration, execution efficiency and self-verification, mean over the 12 PrestoSQL/Trino tasks; one run per task, every model in the same Tencent CodeBuddy agent scaffold). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | CodeBuddy + Gemini 3.5 Flash | 77.7 |
| 2 | CodeBuddy + DeepSeek V4 Pro | 77.7 |
| 3 | CodeBuddy + Kimi K2.6 | 76.2 |
| 4 | CodeBuddy + Kimi K2.7 | 76.1 |
| 5 | CodeBuddy + GLM 5.1 | 75.9 |
| 6 | CodeBuddy + MiniMax M3 | 75.4 |
| 7 | CodeBuddy + GPT 5.5 | 75.3 |
| 8 | CodeBuddy + Claude Opus 4.8 | 74.7 |
| 9 | CodeBuddy + DeepSeek V4 Flash | 74.4 |
| 10 | CodeBuddy + Gemini 3.1 Pro | 72.8 |
Interactive version: theaggregate.ai/benchmark?slug=dataclaweval-prestosql-trino · How It Works · Data refreshed daily, snapshot 2026-09-29.