CFMME - Recognition, Detection and Extraction: leaderboard

Metric: Average (%) of seal recognition 1-NED, table recognition TEDS, formula recognition CDM, seal detection mAP and information extraction F1 on CFMME (6,052 Chinese financial multimodal instances over eight financial image types), zero-shot, mean of three runs, the non-reasoning model of each family; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B Instruct77.18
2GLM-4.5V (Non-reasoning)71.77
3Qwen 3 VL 8B Instruct71.04
4GPT-4o57.89
5ERNIE 4.5 VL 28B A3B56.61
6Llama 3.2 11B Instruct27.42

Interactive version: theaggregate.ai/benchmark?slug=cfmme-recognition-detection-and-extraction · How It Works · Data refreshed daily, snapshot 2026-10-07.