RenoBench - Given Names: leaderboard

Metric: Recall (%) of gold JATS given-names annotations reproduced by the model's markup over 10,000 real-world references (lenient: fields inside valid tags still count when the outer XML is malformed); two-shot prompt; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 2.5 7B Instruct74.5#846
2Qwen 3 8B70.9#667
3Qwen 3 32B69.2#424
4Mistral 7B Instruct (v0.2)69#1310
5GPT-OSS-20B67#499
6Qwen 2.5 3B Instruct64.1#1138
7Qwen 3 0.6B62.7#1440
8Llama 3.1 8B Instruct58.6#1018
9Gemma 3 1B (IT)52.5#1484

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=renobench-given-names · How It Works · Data refreshed daily, snapshot 2026-10-11.