InternVL3.5-14B-Instruct: benchmark results
Provider: Shanghai AI Lab. Access: Open.
Unified ELO 1530 ± 25, rank #651 of 1605 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SiT-Bench - Global Perception & Mapping | 23.64 | Accuracy (%; item-weighted over the category's subtasks) | 76.7 |
| SiT-Bench - Logic & Anomaly Detection | 37.56 | Accuracy (%; item-weighted over the category's subtasks) | 76.7 |
| SiT-Bench - Navigation & Planning | 49.33 | Accuracy (%; item-weighted over the category's subtasks) | 76.7 |
| MCSBench - L3 | 26.72 | Accuracy (%) | 76.3 |
| SiT-Bench | 41.47 | Accuracy (%; item-weighted over the 17 subtasks) | 69 |
| SiT-Bench - Embodied & Fine-grained Perception | 42.81 | Accuracy (%; item-weighted over the category's subtasks) | 63.3 |
| SiT-Bench - Multi-View & Geometric Reasoning | 46.17 | Accuracy (%; item-weighted over the category's subtasks) | 61.7 |
| MCSBench - L2 | 53.27 | Accuracy (%) | 60.5 |
| MCSBench - Overall | 48.41 | Accuracy (%) | 59.2 |
| MCSBench - L1 | 56.32 | Accuracy (%) | 48 |
| K-MetBench | 47.9 | Accuracy (self-reported) | 32.8 |
| Video-IFBench - Selection | 12.5 | Task-gated instruction satisfaction rate (%; TISR on selecti | 29.7 |
Interactive version: theaggregate.ai/model?slug=internvl3-5-14b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-26.