Keye-VL-1.5-8B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1484 ± 1, rank #1099 of 2032 rated models, from 15 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
TempCompass75.5Avg All (%)100
Open Portuguese LLM - BLUEX65.92Accuracy (%)87.7
Open Portuguese LLM - FaQuAD NLI79.18Macro F1 (%)86.6
Open Portuguese LLM - ENEM73.2Accuracy (%)84.5
Open Portuguese LLM - OAB Exams51.94Accuracy (%)81
Open Portuguese LLM - ASSIN2 RTE92.5Macro F1 (%)79.8
Molmo2-VideoCount27.2Exact-count accuracy (%; 533 validation queries)47.5
Molmo2-CapTest25.4Caption F1 (%; LLM-judged statement precision and recall aga45
VI-Bench (Video Prompt Inversion) - Medium (Style and Camera)0.53Inversion Score (0-1; mean of the Prompt Score, GPT-4o (gpt-43.8
VI-Bench (Video Prompt Inversion) - Easy (Semantic Grounding)0.7Inversion Score (0-1; mean of the Prompt Score, GPT-4o (gpt-37.5
VidOmni-Bench9.6F1 (%; sentence-wise event verification: for each of the den33.3
Molmo2 Human Preference - Captioning957Elo rating (bootstrap median, Bradley-Terry; pairwise human 25

Interactive version: theaggregate.ai/model?slug=keye-vl-1-5-8b · How It Works · Data refreshed daily, snapshot 2026-09-26.