SciFigPlag-Bench - Pairwise Detection: leaderboard

Metric: Accuracy (%; yes or no: does the suspicious figure reuse content of the source figure; SciFigPlag-Bench pairs of source and suspicious scientific figures (2,582 plagiarized pairs from documented real cases and taxonomy-guided synthesis, 2,541 visually similar negatives); the same task-specific prompts for every model, temperature 0). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.696.3
2Gemma 4 31B (IT)92.5
3Gemini 3 Flash91.4
4Qwen 3.5 35B A3B87.3
5Gemma 4 26B A4B (IT)87.2
6Qwen 3.6 35B A3B87
7GPT-5.484.4
8Mistral Small 481
9Qwen 3.5 2B75.3
10InternVL3.5-8B74.4
11Gemma 3 4B (IT)67.4
12Pixtral-12B-240952.1

Interactive version: theaggregate.ai/benchmark?slug=scifigplag-bench-pairwise-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.