FinED-Bench (English): leaderboard

Metric: F1 (%; FinED-Bench-EN, 56 English-language financial documents with annotated general-knowledge, domain-knowledge and reasoning errors; a predicted error counts as correct when the extracted sentence matches or contains the annotated sentence and the error type is classified correctly; the paper's error-detection prompt over each document, reasoning models with thinking on unless marked no thinking). Source: arxiv.org. Saturation forecast: Around February 2027. 9 models tracked.

Top models

#ModelScore
1GPT-553.38
2Qwen 3 14B44.52
3Qwen 3 8B38.64
4DeepSeek R1 0528 Qwen3 8B33.94
5Qwen 3 14B (Non-reasoning)23.6
6GPT-4o Mini21.2
7Qwen 3 8B (Non-reasoning)17.58
8Llama 3.1 8B12.9

Interactive version: theaggregate.ai/benchmark?slug=fined-bench-english · How It Works · Data refreshed daily, snapshot 2026-09-26.