DataClawEval - PrestoSQL-Trino: leaderboard

Metric: Rule-based task score (%; per task a weighted sum, usually 0.7 and 0.3, of the artifact score, checked row by row against live databases by task-specific deterministic scripts, and the process score for exploration, execution efficiency and self-verification, mean over the 12 PrestoSQL/Trino tasks; one run per task, every model in the same Tencent CodeBuddy agent scaffold). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 16 models tracked.

Top models

#ModelScore
1CodeBuddy + Gemini 3.5 Flash77.7
2CodeBuddy + DeepSeek V4 Pro77.7
3CodeBuddy + Kimi K2.676.2
4CodeBuddy + Kimi K2.776.1
5CodeBuddy + GLM 5.175.9
6CodeBuddy + MiniMax M375.4
7CodeBuddy + GPT 5.575.3
8CodeBuddy + Claude Opus 4.874.7
9CodeBuddy + DeepSeek V4 Flash74.4
10CodeBuddy + Gemini 3.1 Pro72.8

Interactive version: theaggregate.ai/benchmark?slug=dataclaweval-prestosql-trino · How It Works · Data refreshed daily, snapshot 2026-09-29.