PaliBench - COMET: leaderboard
Metric: COMET (0-100): mean COMET score over the three references, times 100, 1,700 Pali canon passages translated into English, each scored against three human reference translations (Bodhi, Sujato, Thanissaro), corpus level, models through OpenRouter; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 73.1 |
| 2 | Gemini 3 Pro (Preview) | 72.9 |
| 3 | Claude Opus 4.5 | 72.4 |
| 4 | DeepSeek V3.2 | 71.8 |
| 5 | GPT-5.2 | 70.9 |
| 6 | Kimi K2.5 | 70.7 |
| 7 | Qwen 3 235B A22B 2507 Instruct | 70.6 |
| 8 | GLM-4.7 | 70.3 |
| 9 | Grok 4.1 Fast | 69.6 |
| 10 | Llama 3.3 70B Instruct | 68.1 |
Interactive version: theaggregate.ai/benchmark?slug=palibench-comet · How It Works · Data refreshed daily, snapshot 2026-10-07.