BaFCo - Coarse Layout Analysis (CoT): leaderboard
Metric: Mean average precision at IoU 0.3 (0-100): class-wise average precision of predicted layout boxes, greedily matched to ground truth at IoU 0.3, averaged over the 5 coarse entity classes, x100; 200 Bangladeshi government forms (316 pages) with 16,382 annotated entities; chain-of-thought prompt; reasoning_effort low or high as labelled; schema-invalid outputs excluded. Source: arxiv.org. Saturation forecast: Around 2029. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (High) | 24.44 |
| 2 | GPT-5.2 (Low) | 16.71 |
| 3 | Kimi K2.5 (High) | 13.61 |
| 4 | GPT-5.2 (High) | 13.54 |
| 5 | Claude Opus 4.6 (High) | 3.58 |
| 6 | Claude Opus 4.6 (Low) | 1.31 |
Interactive version: theaggregate.ai/benchmark?slug=bafco-coarse-layout-analysis-cot · How It Works · Data refreshed daily, snapshot 2026-09-29.