Grip on LLMs - Dutch tinyTruthfulQA: leaderboard

Metric: Accuracy (%; GP-IRT estimate from 100 tinyBenchmarks items, Dutch machine translation). Source: arxiv.org. 31 models tracked.

Top models

#ModelScore
1Mistral Small 378
2GPT-5 Mini77
3Mistral Medium 376
4GPT-575
5Qwen 3 32B71
6DeepSeek R1 Distill Qwen 32B70
7GPT-4o69
8GPT-5 Nano64
9Gemma 3 27B (IT)63
10Mistral Large 363
11Gemma 3 12B (IT)62
12GPT-OSS-20B62
13GPT-4o Mini61
14aya-expanse-32B61
15Qwen 3 8B59

Interactive version: theaggregate.ai/benchmark?slug=grip-on-llms-dutch-tinytruthfulqa · How It Works · Data refreshed daily, snapshot 2026-09-19.