JavaVulBench: leaderboard
Metric: Detection F1 (%; vulnerable methods as the positive class) of zero-shot method-level Java vulnerability classification on the fixed 200-sample stratified probe (seed 42) of the project-disjoint test split, API-served LLMs through OpenRouter; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 41.9 |
| 2 | Claude Sonnet 4 | 41.6 |
| 3 | GPT-4.1 Mini | 40 |
| 4 | DeepSeek V3 | 30 |
Interactive version: theaggregate.ai/benchmark?slug=javavulbench · How It Works · Data refreshed daily, snapshot 2026-09-29.