PaliBench - COMET: leaderboard

Metric: COMET (0-100): mean COMET score over the three references, times 100, 1,700 Pali canon passages translated into English, each scored against three human reference translations (Bodhi, Sujato, Thanissaro), corpus level, models through OpenRouter; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 10 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)73.1
2Gemini 3 Pro (Preview)72.9
3Claude Opus 4.572.4
4DeepSeek V3.271.8
5GPT-5.270.9
6Kimi K2.570.7
7Qwen 3 235B A22B 2507 Instruct70.6
8GLM-4.770.3
9Grok 4.1 Fast69.6
10Llama 3.3 70B Instruct68.1

Interactive version: theaggregate.ai/benchmark?slug=palibench-comet · How It Works · Data refreshed daily, snapshot 2026-10-07.