RoundTripCodeEval - LZW: leaderboard

Metric: Mean of exact match, edit similarity and pass@5 (%) over the four round-trip tasks for LZW (output prediction and input prediction, each directly and through the inverted function), 250 inputs, zero-shot with one worked example, five completions. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1QwQ-32B24.14
2DeepSeek R1 Distill Qwen 32B23.81
3Qwen 2.5 Coder 32B Instruct21.06
4DeepSeek R1 Distill Qwen 14B14.03
5Codestral-22B-v0.17.77
6Phi-3.5-mini-instruct4.81
7Qwen 2.5 7B Instruct4.46
8Yi-Coder-9B-Chat4.21
9Phi-3 Mini 128K Instruct3.65
10Mistral 7B Instruct (v0.3)2.25
11Llama 3.1 8B Instruct0.82
12DeepSeek R1 Distill Qwen 1.5B0.32
13Llama 3.2 1B Instruct0.05

Interactive version: theaggregate.ai/benchmark?slug=roundtripcodeeval-lzw · How It Works · Data refreshed daily, snapshot 2026-09-26.