RedCode — leaderboard

Tests risky code execution and harmful code generation capabilities. Evaluates both ability to generate exploits (RedCode-Gen) and willingness to execute dangerous code (RedCode-Exec).

Metric: Score (%). Source: huggingface.co. Status: saturated. 15 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct75.54
2GPT-4o73.67
3Claude 3.5 Sonnet65.15
4GPT-464.14
5Mistral 7B61.7
6Llama 3.1 8B Instruct61.35
7GPT-3.554.29
8Llama 3 8B Instruct52.38
9Llama 2 7B44.38
10Claude 3 Opus2.2

Interactive version: theaggregate.ai/benchmark?slug=redcode · How the rankings work · Data refreshed daily, snapshot 2026-07-22.