DataClawEval: leaderboard
Metric: Rule-based task score (%; per task a weighted sum, usually 0.7 and 0.3, of the artifact score, checked row by row against live databases by task-specific deterministic scripts, and the process score for exploration, execution efficiency and self-verification, mean over all 100 tasks; one run per task, every model in the same Tencent CodeBuddy agent scaffold). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | CodeBuddy + GPT 5.5 | 74.9 |
| 2 | CodeBuddy + Claude Opus 4.8 | 74.3 |
| 3 | CodeBuddy + Claude Sonnet 5 | 73.8 |
| 4 | CodeBuddy + Gemini 3.1 Pro | 73.7 |
| 5 | CodeBuddy + Gemini 3.5 Flash | 73.3 |
| 6 | CodeBuddy + DeepSeek V4 Flash | 73 |
| 7 | CodeBuddy + MiniMax M3 | 71.8 |
| 8 | CodeBuddy + GLM 5.1 | 71.6 |
| 9 | CodeBuddy + DeepSeek V4 Pro | 70.6 |
| 10 | CodeBuddy + Kimi K2.6 | 69 |
Interactive version: theaggregate.ai/benchmark?slug=dataclaweval · How It Works · Data refreshed daily, snapshot 2026-09-29.