VIVID - Linguistic and Cultural Classification: leaderboard

Metric: Accuracy (%; exact-match labelling of the linguistic and cultural complexity traits of each idiom or proverb, three-shot, LM Evaluation Harness). Source: arxiv.org. Saturation forecast: Around May 2027. 8 models tracked.

Top models

#ModelScore
1GPT-4o18.4
2Gemini 2.5 Flash14.4
3Llama 4 Scout Base4.2
4Qwen 3 14B3.1

Interactive version: theaggregate.ai/benchmark?slug=vivid-linguistic-and-cultural-classification · How It Works · Data refreshed daily, snapshot 2026-09-26.