Keye-VL-1.5-8B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1484 ± 1, rank #1099 of 2032 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| TempCompass | 75.5 | Avg All (%) | 100 |
| Open Portuguese LLM - BLUEX | 65.92 | Accuracy (%) | 87.7 |
| Open Portuguese LLM - FaQuAD NLI | 79.18 | Macro F1 (%) | 86.6 |
| Open Portuguese LLM - ENEM | 73.2 | Accuracy (%) | 84.5 |
| Open Portuguese LLM - OAB Exams | 51.94 | Accuracy (%) | 81 |
| Open Portuguese LLM - ASSIN2 RTE | 92.5 | Macro F1 (%) | 79.8 |
| Molmo2-VideoCount | 27.2 | Exact-count accuracy (%; 533 validation queries) | 47.5 |
| Molmo2-CapTest | 25.4 | Caption F1 (%; LLM-judged statement precision and recall aga | 45 |
| VI-Bench (Video Prompt Inversion) - Medium (Style and Camera) | 0.53 | Inversion Score (0-1; mean of the Prompt Score, GPT-4o (gpt- | 43.8 |
| VI-Bench (Video Prompt Inversion) - Easy (Semantic Grounding) | 0.7 | Inversion Score (0-1; mean of the Prompt Score, GPT-4o (gpt- | 37.5 |
| VidOmni-Bench | 9.6 | F1 (%; sentence-wise event verification: for each of the den | 33.3 |
| Molmo2 Human Preference - Captioning | 957 | Elo rating (bootstrap median, Bradley-Terry; pairwise human | 25 |
Interactive version: theaggregate.ai/model?slug=keye-vl-1-5-8b · How It Works · Data refreshed daily, snapshot 2026-09-26.