CHaystack - Gold-Image QA (F1): leaderboard
Metric: Character-level F1 (%; 4,543 Chinese document questions (academic papers, web pages, advertisements, photographed documents) answered from the gold evidence image, answers normalized for punctuation, whitespace and case; at most 128 generated tokens). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen2.5-VL-3B-Instruct | 61.62 |
| 2 | InternVL2.5-4B | 42.64 |
| 3 | LLaVA-OneVision-Qwen2-0.5B | 19.68 |
Interactive version: theaggregate.ai/benchmark?slug=chaystack-gold-image-qa-f1 · How It Works · Data refreshed daily, snapshot 2026-09-29.