ImplicitBBQ - Disambiguated: leaderboard
Metric: Accuracy (%, times 100) in disambiguated implicit contexts, where the context names which person the question is about and the answer is that person: demographic identity is conveyed only by human-validated characteristic cues (kappa at least 0.60) instead of labels, about 8k implicit instances built from 760 BBQ and BharatBBQ contexts and 60 cues across age, gender, region, religion, caste and socioeconomic status, negative and non-negative questions, zero-shot single-word answers, model-wise average over demographics; open models on the full set, closed models on a balanced 600-sample subset; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 100 |
| 2 | Claude Haiku 4.5 | 99 |
| 3 | GPT-OSS-120B | 98 |
| 4 | Qwen 3 32B | 98 |
| 5 | Gemini 3.1 Flash Lite | 98 |
| 6 | Llama 3.3 70B Instruct | 97 |
| 7 | GPT-OSS-20B | 97 |
| 8 | Llama 3.1 8B Instruct | 96 |
| 9 | Mistral 7B Instruct | 95 |
| 10 | Qwen 2.5 7B Instruct | 94 |
| 11 | Phi-4 Mini Instruct | 93 |
Interactive version: theaggregate.ai/benchmark?slug=implicitbbq-disambiguated · How It Works · Data refreshed daily, snapshot 2026-10-07.