DataClawEval - Chinese Tasks: leaderboard
Metric: Rule-based task score (%; per task a weighted sum, usually 0.7 and 0.3, of the artifact score, checked row by row against live databases by task-specific deterministic scripts, and the process score for exploration, execution efficiency and self-verification, mean over the 50 Chinese-language tasks; one run per task, every model in the same Tencent CodeBuddy agent scaffold). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | CodeBuddy + Claude Sonnet 5 | 74.1 |
| 2 | CodeBuddy + Claude Opus 4.8 | 73.6 |
| 3 | CodeBuddy + GPT 5.5 | 72.9 |
| 4 | CodeBuddy + Gemini 3.5 Flash | 72.3 |
| 5 | CodeBuddy + Gemini 3.1 Pro | 72.1 |
| 6 | CodeBuddy + DeepSeek V4 Flash | 69.7 |
| 7 | CodeBuddy + MiniMax M3 | 69.4 |
| 8 | CodeBuddy + Kimi K2.6 | 68.3 |
| 9 | CodeBuddy + DeepSeek V4 Pro | 66.9 |
| 10 | CodeBuddy + GLM 5.1 | 66.6 |
Interactive version: theaggregate.ai/benchmark?slug=dataclaweval-chinese-tasks · How It Works · Data refreshed daily, snapshot 2026-09-29.