VSysBench: leaderboard

Metric: Joint satisfaction rate (%; per-sample task score counted only when the judged constraint score reaches 0.8, averaged over the aligned and misaligned halves; VSysBench: 2,258 verified MM-Vet v2 image questions, each under a system-message constraint from 22 sub-categories in 5 categories, once with an aligned and once with a conflicting user message (4,516 instances); default inference settings; GPT-5-mini judge at temperature 0). Source: arxiv.org. Saturation forecast: Around November 2027. 16 models tracked.

Top models

#ModelScore
1GPT-5.436.2
2Claude Opus 4.736.2
3Claude Sonnet 4.631.9
4GPT-5.4 Mini28.3
5Claude Haiku 4.522
6GPT-5.4 Nano19.2
7GPT-4o17.6
8Qwen 3 VL 32B16.6
9InternVL3.5-8B12.7
10Phi-4 Multimodal Instruct4.4

Interactive version: theaggregate.ai/benchmark?slug=vsysbench · How It Works · Data refreshed daily, snapshot 2026-09-29.