Pseudo2CodeQA: leaderboard

Metric: Overall rubric score (1-5) from a GPT-5 judge over correctness, completeness, relevance, clarity, reasoning and pseudocode adherence with execution testing, weighting correctness, adherence and execution; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro4.31
2GPT-3.5 Turbo4.06
3CodeLlama-7B-hf2.99
4DeepSeek R1 Distill Llama 8B2.97
5DeepSeek-R1-Distill-Qwen-7B2.95
6Phi-22.95
7gemma-7B2.94
8Mistral 7B2.94
9Llama 3.2 3B2.44
10gemma-2B2.16
11DeepSeek R1 Distill Qwen 1.5B2.1

Interactive version: theaggregate.ai/benchmark?slug=pseudo2codeqa · How It Works · Data refreshed daily, snapshot 2026-09-26.