CyberSecEval2 Prompt Injection — leaderboard

Metric: Accuracy (%). Source: github.com. 10 models tracked.

Top models

#ModelScore
1DeepSeek V3 Chat48.01
2Llama 3.2 90B Vision Instruct37.35
3Llama 3.3 70B Instruct36.45
4Grok 2 (1212)29.18
5mistral-large-latest28.88
6mistral-small-latest28.59
7GPT-4o Mini (2024-07-18)25.8
8Gemini 2.0 Flash (001)24.7
9Claude 3.7 Sonnet (20250219)17.43
10GPT-4o (2024-08-06)17.13

Interactive version: theaggregate.ai/benchmark?slug=cyberseceval2-prompt-injection · How the rankings work · Data refreshed daily, snapshot 2026-07-22.