Pseudo2CodeQA: leaderboard
Metric: Overall rubric score (1-5) from a GPT-5 judge over correctness, completeness, relevance, clarity, reasoning and pseudocode adherence with execution testing, weighting correctness, adherence and execution; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 4.31 |
| 2 | GPT-3.5 Turbo | 4.06 |
| 3 | CodeLlama-7B-hf | 2.99 |
| 4 | DeepSeek R1 Distill Llama 8B | 2.97 |
| 5 | DeepSeek-R1-Distill-Qwen-7B | 2.95 |
| 6 | Phi-2 | 2.95 |
| 7 | gemma-7B | 2.94 |
| 8 | Mistral 7B | 2.94 |
| 9 | Llama 3.2 3B | 2.44 |
| 10 | gemma-2B | 2.16 |
| 11 | DeepSeek R1 Distill Qwen 1.5B | 2.1 |
Interactive version: theaggregate.ai/benchmark?slug=pseudo2codeqa · How It Works · Data refreshed daily, snapshot 2026-09-26.