CodeTransOcean - LLMTrans — leaderboard
CodeTransOcean sub-benchmark evaluating LLM-specific code translation capabilities. Evaluated by Debug Success Rate (DSR).
Metric: DSR (%). Source: yuchen814.github.io. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-3.5 (3rd Debug) | 52.57 |
| 2 | GPT-3.5 (One-shot) | 50.29 |
| 3 | GPT-3.5 (Zero-shot) | 48.57 |
| 4 | GPT-3.5 (CoT) | 48.29 |
Interactive version: theaggregate.ai/benchmark?slug=codetransocean-llmtrans · How the rankings work · Data refreshed daily, snapshot 2026-07-22.