SecCodePLT — leaderboard

Evaluates LLMs on secure code generation and cyberattack helpfulness across 13 subtasks covering insecure coding patterns, reconnaissance, weaponization, and discovery.

Metric: Score (%). Source: huggingface.co. 6 models tracked.

Top models

#ModelScore
1CodeLlama-34B-Instruct67.88
2Mixtral 8x22B58
3Claude 3.5 Sonnet54.4
4Llama 3.1 70B44.75
5GPT-4o44.15

Interactive version: theaggregate.ai/benchmark?slug=seccodeplt · How the rankings work · Data refreshed daily, snapshot 2026-07-22.