OVIBench (Open-Ended) - Intent Fulfillment: leaderboard
Metric: Intent fulfillment accuracy (%; share of post-interruption responses that satisfy the user actual intent, judged 0/1 by Qwen3-VL-235B-A22B; open-ended subset of 13,145 samples; offline simulation of an interruption arriving while the model is generating its answer about a streaming video (3,200 videos from ActivityNet, MovieChat, QVHighlights, UCF-Crime and YouCook2)). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 VL 8B | 81.07 |
| 2 | Seed-1.6 | 79.42 |
| 3 | Gemini 2.5 Flash | 64.58 |
| 4 | Qwen 2.5 VL 7B | 50.11 |
Interactive version: theaggregate.ai/benchmark?slug=ovibench-open-ended-intent-fulfillment · How It Works · Data refreshed daily, snapshot 2026-09-26.