BaFCo - Coarse Layout Analysis (Zero-shot): leaderboard

Metric: Mean average precision at IoU 0.3 (0-100): class-wise average precision of predicted layout boxes, greedily matched to ground truth at IoU 0.3, averaged over the 5 coarse entity classes, x100; 200 Bangladeshi government forms (316 pages) with 16,382 annotated entities; zero-shot prompt; reasoning_effort low or high as labelled; schema-invalid outputs excluded. Source: arxiv.org. Saturation forecast: Around 2029. 10 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (High)26.46
2GPT-5.2 (Low)18
3GPT-5.2 (High)14.34
4Kimi K2.5 (High)13.36
5Claude Opus 4.6 (High)4.4
6Claude Opus 4.6 (Low)1.68

Interactive version: theaggregate.ai/benchmark?slug=bafco-coarse-layout-analysis-zero-shot · How It Works · Data refreshed daily, snapshot 2026-09-29.