RoundTripCodeEval - Huffman Coding: leaderboard

Metric: Mean of exact match, edit similarity and pass@5 (%) over the four round-trip tasks for Huffman coding (output prediction and input prediction, each directly and through the inverted function), 250 inputs, zero-shot with one worked example, five completions. Source: arxiv.org. Saturation forecast: Around 2029. 15 models tracked.

Top models

#ModelScore
1QwQ-32B5.5
2DeepSeek R1 Distill Qwen 32B3.98
3DeepSeek R1 Distill Qwen 14B3.15
4Qwen 2.5 Coder 32B Instruct3.15
5Qwen 2.5 7B Instruct2.65
6Yi-Coder-9B-Chat1.87
7Phi-3.5-mini-instruct1.7
8Phi-3 Mini 128K Instruct1.54
9Codestral-22B-v0.11.5
10Llama 3.1 8B Instruct1.13
11Mistral 7B Instruct (v0.3)0.99
12DeepSeek R1 Distill Qwen 1.5B0.64
13Llama 3.2 1B Instruct0.08

Interactive version: theaggregate.ai/benchmark?slug=roundtripcodeeval-huffman-coding · How It Works · Data refreshed daily, snapshot 2026-09-26.