CFMME - Information Extraction: leaderboard

Metric: Field-level F1 (%) of key information extracted from financial documents on CFMME (6,052 Chinese financial multimodal instances over eight financial image types), zero-shot, mean of three runs, the non-reasoning model of each family; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B Instruct93.43
2GLM-4.5V (Non-reasoning)89.03
3Qwen 3 VL 8B Instruct88.86
4ERNIE 4.5 VL 28B A3B80.71
5GPT-4o66.88
6Llama 3.2 11B Instruct47.52

Interactive version: theaggregate.ai/benchmark?slug=cfmme-information-extraction · How It Works · Data refreshed daily, snapshot 2026-10-07.