RoundTripCodeEval - Run-Length Encoding: leaderboard

Metric: Mean of exact match, edit similarity and pass@5 (%) over the four round-trip tasks for run-length encoding (output prediction and input prediction, each directly and through the inverted function), 250 inputs, zero-shot with one worked example, five completions. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1QwQ-32B57.23
2Qwen 2.5 Coder 32B Instruct41.51
3DeepSeek R1 Distill Qwen 32B36.37
4Codestral-22B-v0.130.68
5DeepSeek R1 Distill Qwen 14B26.97
6Qwen 2.5 7B Instruct17.39
7Phi-3.5-mini-instruct13.67
8Phi-3 Mini 128K Instruct12.01
9Yi-Coder-9B-Chat11.85
10Mistral 7B Instruct (v0.3)11.52
11Llama 3.1 8B Instruct9.36
12DeepSeek R1 Distill Qwen 1.5B4.3
13Llama 3.2 1B Instruct0.15

Interactive version: theaggregate.ai/benchmark?slug=roundtripcodeeval-run-length-encoding · How It Works · Data refreshed daily, snapshot 2026-09-26.