StereoSet: leaderboard
Stereotype bias benchmark for language models, testing model preference among stereotypical, anti-stereotypical, and unrelated continuations across social domains.
Metric: ICAT Score. Source: huggingface.co. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-2 | 72.97 |
| 2 | GPT-2 Medium | 71.73 |
| 3 | GPT-2 Large | 70.54 |
| 4 | text-davinci-002 | 60.8 |
Interactive version: theaggregate.ai/benchmark?slug=stereoset · How It Works · Data refreshed daily, snapshot 2026-09-05.