Synthetic Users Benchmark - WVS Individual Accuracy: leaderboard
Metric: Individual accuracy (%; World Values Survey Wave 7, 16 WorldValuesBench value questions over 63 countries, 1,458 sampled respondent-question rows, profile of age group, sex, education, settlement type and country; single-answer prompt (Style A): given the demographic profile of a real respondent, return one answer option; exact-match share of sampled respondent-question pairs whose actual answer is predicted; two sampling runs per cell at temperature 1.0 with a 100-token cap, unreadable answers flagged invalid). Source: arxiv.org. Saturation forecast: Around 2034. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.3 70B Instruct | 24.9 |
| 2 | Claude Sonnet 4.6 | 23.6 |
| 3 | Llama 3.1 8B Instruct | 17.9 |
| 4 | Claude Haiku 4.5 | 17 |
Interactive version: theaggregate.ai/benchmark?slug=synthetic-users-benchmark-wvs-individual-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.