JavaVulBench: leaderboard

Metric: Detection F1 (%; vulnerable methods as the positive class) of zero-shot method-level Java vulnerability classification on the fixed 200-sample stratified probe (seed 42) of the project-disjoint test split, API-served LLMs through OpenRouter; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1GPT-4o41.9
2Claude Sonnet 441.6
3GPT-4.1 Mini40
4DeepSeek V330

Interactive version: theaggregate.ai/benchmark?slug=javavulbench · How It Works · Data refreshed daily, snapshot 2026-09-29.