PaliBench - chrF++: leaderboard

Metric: chrF++ (0-100), character n-gram F-score with word bigrams (sacrebleu), 1,700 Pali canon passages translated into English, each scored against three human reference translations (Bodhi, Sujato, Thanissaro), corpus level, models through OpenRouter; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 10 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)68.5
2Gemini 3 Flash (Preview)65.9
3Claude Opus 4.565.6
4DeepSeek V3.264.1
5Kimi K2.561.2
6GPT-5.259.1
7GLM-4.757.1
8Qwen 3 235B A22B 2507 Instruct55.7
9Grok 4.1 Fast54.8
10Llama 3.3 70B Instruct48.1

Interactive version: theaggregate.ai/benchmark?slug=palibench-chrf-plus-plus · How It Works · Data refreshed daily, snapshot 2026-10-07.