BFI-Adapt: leaderboard

Metric: BFI-Adapt composite (%; mean over the 27 event-trait pairs with a definite human change direction of max(0, weighted kappa) x (2 DCR - 1) x an indicator that the dominant item-change direction matches the human prior; 100 personas simulate 11 life events and answer the BFI-44 before and after; non-reasoning mode). Source: arxiv.org. Saturation forecast: Around 2029. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Minimal)34.8
2GLM-4.6 (Non-reasoning)32.2
3Qwen 3 235B A22B (Non-reasoning)28.7
4Claude Haiku 4.522.5
5Kimi K2 090522.4
6Seed 2.0 Pro (Non-reasoning)21.8
7Claude Sonnet 4.621.2
8Qwen 3.5 9B (Non-reasoning)19.4
9GPT-4.1 Mini19.2
10GPT-5.317.8
11DeepSeek V4 Pro (Non-reasoning)17.6
12MiMo-V2.5-Pro (Non-reasoning)7.1
13Qwen 2.5 14B Instruct4.9
14InternLM3-8B-Instruct4.4

Interactive version: theaggregate.ai/benchmark?slug=bfi-adapt · How It Works · Data refreshed daily, snapshot 2026-09-26.