RoundTripCodeEval - Arithmetic Coding: leaderboard

Metric: Mean of exact match, edit similarity and pass@5 (%) over the four round-trip tasks for arithmetic coding (output prediction and input prediction, each directly and through the inverted function), 250 inputs, zero-shot with one worked example, five completions. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1QwQ-32B15.71
2DeepSeek R1 Distill Qwen 32B12.74
3DeepSeek R1 Distill Qwen 14B10.08
4Qwen 2.5 Coder 32B Instruct8.45
5Qwen 2.5 7B Instruct6.55
6Llama 3.1 8B Instruct4.45
7Phi-3.5-mini-instruct2.85
8Phi-3 Mini 128K Instruct2.6
9Yi-Coder-9B-Chat2.2
10Mistral 7B Instruct (v0.3)2.05
11DeepSeek R1 Distill Qwen 1.5B1.83
12Codestral-22B-v0.11.76
13Llama 3.2 1B Instruct0.34

Interactive version: theaggregate.ai/benchmark?slug=roundtripcodeeval-arithmetic-coding · How It Works · Data refreshed daily, snapshot 2026-09-26.