ImplicitBBQ - Disambiguated: leaderboard

Metric: Accuracy (%, times 100) in disambiguated implicit contexts, where the context names which person the question is about and the answer is that person: demographic identity is conveyed only by human-validated characteristic cues (kappa at least 0.60) instead of labels, about 8k implicit instances built from 760 BBQ and BharatBBQ contexts and 60 cues across age, gender, region, religion, caste and socioeconomic status, negative and non-negative questions, zero-shot single-word answers, model-wise average over demographics; open models on the full set, closed models on a balanced 600-sample subset; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1GPT-5 Mini100
2Claude Haiku 4.599
3GPT-OSS-120B98
4Qwen 3 32B98
5Gemini 3.1 Flash Lite98
6Llama 3.3 70B Instruct97
7GPT-OSS-20B97
8Llama 3.1 8B Instruct96
9Mistral 7B Instruct95
10Qwen 2.5 7B Instruct94
11Phi-4 Mini Instruct93

Interactive version: theaggregate.ai/benchmark?slug=implicitbbq-disambiguated · How It Works · Data refreshed daily, snapshot 2026-10-07.