StereoSet — leaderboard
Stereotype bias benchmark for language models, testing model preference among stereotypical, anti-stereotypical, and unrelated continuations across social domains.
Metric: ICAT Score. Source: huggingface.co. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-2 | 72.97 |
| 2 | GPT-2 Medium | 71.73 |
Interactive version: theaggregate.ai/benchmark?slug=stereoset · How the rankings work · Data refreshed daily, snapshot 2026-07-22.