RVDB: leaderboard
Metric: Average decision flip rate (%) over the ten pairwise role-gender configuration comparisons: share of decision units whose supported decision changes when only the gender of the two roles changes (position-balanced, zero-shot, temperature 0); lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 Max | 3.37 |
| 2 | GLM-4 9B Chat | 4.89 |
| 3 | Qwen 2.5 7B Instruct | 5.37 |
| 4 | Qwen 3 14B | 5.64 |
| 5 | GPT-4o Mini | 6.22 |
| 6 | Llama 3.1 8B Instruct | 9.5 |
| 7 | Llama 3 8B Instruct | 9.61 |
Interactive version: theaggregate.ai/benchmark?slug=rvdb · How It Works · Data refreshed daily, snapshot 2026-09-29.