Khondo (MonoSeq): leaderboard
Metric: Packet score (-0.5 to 1; mean of the clustering score, the average of V-measure and Rand index, and the ordering score, mean Kendall tau-a over predicted multi-page documents; forms of one domain concatenated in order; 210 test packets of 5 to 20 page images of Bangladeshi government forms (Bangla and English, 14 domains); zero-shot order-aware prompt, all page images in one request, JSON output, temperature 1.0, reasoning disabled, invalid outputs retried up to three times). Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.5 Flash (Minimal) | 0.9 |
| 2 | Qwen 3.6 Plus (Non-reasoning) | 0.9 |
| 3 | GPT-5.4 (Non-reasoning) | 0.83 |
| 4 | GLM-4.6V (Non-reasoning) | 0.77 |
| 5 | Kimi K2.5 (Non-reasoning) | 0.76 |
Interactive version: theaggregate.ai/benchmark?slug=khondo-monoseq · How It Works · Data refreshed daily, snapshot 2026-09-29.