RVDB: leaderboard

Metric: Average decision flip rate (%) over the ten pairwise role-gender configuration comparisons: share of decision units whose supported decision changes when only the gender of the two roles changes (position-balanced, zero-shot, temperature 0); lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Qwen 3 Max3.37
2GLM-4 9B Chat4.89
3Qwen 2.5 7B Instruct5.37
4Qwen 3 14B5.64
5GPT-4o Mini6.22
6Llama 3.1 8B Instruct9.5
7Llama 3 8B Instruct9.61

Interactive version: theaggregate.ai/benchmark?slug=rvdb · How It Works · Data refreshed daily, snapshot 2026-09-29.