InternVL3.5-8B: benchmark results

OpenGVLab's 8B InternVL3.5 vision-language model (August 2025) with a Qwen3 backbone and Cascade RL training for multimodal reasoning. Provider: Shanghai AI Lab. Released 2025-08-26. Access: Open.

Unified ELO 1556 ± 4, rank #581 of 1607 rated models, from 723 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ArtECulture - Chinese54.34Accuracy (%; Chinese-speaking culture labels; zero-shot pred100
CapRiCorn-1K-V - Referential Consistency5.4Subject referential consistency (%; share of same-subject de100
Holtercare-Bench - Diagnosis (Video)42.13Accuracy (%; multiple-choice closed QA on the overall diagno100
Holtercare-Bench - Presence (Video)67.97Accuracy (%; multiple-choice closed QA on whether a given rh100
MMBU - Grounded Classification (Boxes)43.9Closed-ended classification of a region marked by a bounding100
MPCI-Bench - Seed Tier Probing97.8Probe Accuracy (%)100
MPCI-Bench - Story Tier Probing95.1Probe Accuracy (%)100
RoboProcessBench - Phase Recognition37.4Accuracy (%) on 1,274 questions asking which coarse process 100
MMBU - Ungrounded Classification (Open-Ended)11.1Open-ended (free-form) ungrounded classification of the whol93.8
RoboProcessBench - Current Primitive Recognition36.8Accuracy (%) on 359 questions asking which low-level primiti92.3
SIS-Bench - Action Recognition56.1Accuracy (%; 686 action recognition questions; four-option m92
AgroVG39.63All@.5 (self-reported)91.7

Interactive version: theaggregate.ai/model?slug=internvl3-5-8b · How It Works · Data refreshed daily, snapshot 2026-09-29.