DataClawEval - MySQL: leaderboard
Metric: Rule-based task score (%; per task a weighted sum, usually 0.7 and 0.3, of the artifact score, checked row by row against live databases by task-specific deterministic scripts, and the process score for exploration, execution efficiency and self-verification, mean over the 20 MySQL tasks; one run per task, every model in the same Tencent CodeBuddy agent scaffold). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | CodeBuddy + GPT 5.5 | 88.8 |
| 2 | CodeBuddy + Gemini 3.5 Flash | 85.2 |
| 3 | CodeBuddy + Claude Opus 4.8 | 85.1 |
| 4 | CodeBuddy + Claude Sonnet 5 | 84.2 |
| 5 | CodeBuddy + DeepSeek V4 Pro | 83.7 |
| 6 | CodeBuddy + Kimi K2.6 | 83.6 |
| 7 | CodeBuddy + GLM 5.2 | 83 |
| 8 | CodeBuddy + Hy3 | 81.9 |
| 9 | CodeBuddy + MiniMax M3 | 81.6 |
| 10 | CodeBuddy + GPT 5.3 Codex | 81.6 |
Interactive version: theaggregate.ai/benchmark?slug=dataclaweval-mysql · How It Works · Data refreshed daily, snapshot 2026-09-29.