CCBench-Health - Overall Norm Adaptation: leaderboard

Metric: Overall norm adaptation rate (%; share of norm-specific checklist recommendations, follow or avoid, that the answer satisfies; 60 personas with graded norm-adherence states across six cultures (Afghan, Burmese, Chinese, Maori, Nepali, Vietnamese), 52 real health questions each (3,120 interactions); GPT-5.2 checklist evaluator; temperature 0.7; the persona conversation history is given before the question and the norms must be inferred from it)). Source: arxiv.org. Saturation forecast: Around 2030. 5 models tracked.

Top models

#ModelScore
1GPT-5.228.7
2Gemini 2.5 Pro24.5
3Qwen 3.5 397B A17B24.2
4Llama 3.2 90B19.8

Interactive version: theaggregate.ai/benchmark?slug=ccbench-health-overall-norm-adaptation · How It Works · Data refreshed daily, snapshot 2026-09-29.