WorkBuddy Bench (CodeBuddy Code) - Security: leaderboard

Metric: Score (0-100) on the 60 red-team and blue-team security tasks scored by programmatic per-task scorers, run in the CodeBuddy Code harness in think mode, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 7 models tracked.

Top models

#ModelScore
1CodeBuddy Code + GLM-5.276.32
2CodeBuddy Code + MiniMax-M374.14
3CodeBuddy Code + DeepSeek-V4-Pro (Thinking)70.04
4CodeBuddy Code + DeepSeek-V4-Flash (Thinking)67.11
5CodeBuddy Code + Hy364.5
6CodeBuddy Code + GPT-5.564.39
7CodeBuddy Code + Claude Opus 4.8 (Thinking)64.37

Interactive version: theaggregate.ai/benchmark?slug=workbuddy-bench-codebuddy-code-security · How It Works · Data refreshed daily, snapshot 2026-09-29.