SciIntegrity-Bench: leaderboard

Metric: Scenarios with explicit misconduct (count of 33; three scenarios in each of 11 trap categories; minimal ReAct agent, temperature 0; human annotation against pre-specified violation criteria). Source: arxiv.org. Saturation forecast: Around June 2027. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.61
2GPT-5.22
3DeepSeek V3.23
4Gemini 3.1 Pro (Preview)5
5Qwen 3.5 397B A17B6
6GLM-57
7Kimi K2.512

Interactive version: theaggregate.ai/benchmark?slug=sciintegrity-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.