IndicQE-APE - Segment-Level Quality Estimation: leaderboard
Metric: Macro Spearman correlation with human direct assessment (-1 to 1; segment-level quality estimation of machine translation by zero-shot GEMBA-DA prompting, the model returning a 0-100 score per segment; within-language correlation averaged over nine directions (six English to Indic, three into English) of the 12,730-segment challenge test; each model over the segments it answers in the required format). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 0.62 |
| 2 | Sarvam M | 0.44 |
| 3 | aya-expanse-32B | 0.4 |
| 4 | Llama 3.2 3B Instruct | 0.22 |
Interactive version: theaggregate.ai/benchmark?slug=indicqe-ape-segment-level-quality-estimation · How It Works · Data refreshed daily, snapshot 2026-09-29.