SocialBias-Bench: leaderboard

Metric: Code bias score (%; share of executable generated snippets whose output changes when any one sensitive attribute changes; 343 human-centred coding tasks, five completions each at temperature 1.0). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1CodeLlama-70B-Instruct-hf28.34
2Claude 3 Haiku (20240307)36.33
3GPT-3.5 Turbo (0125)60.58

Interactive version: theaggregate.ai/benchmark?slug=socialbias-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.