IndicQE-APE - Segment-Level Quality Estimation: leaderboard

Metric: Macro Spearman correlation with human direct assessment (-1 to 1; segment-level quality estimation of machine translation by zero-shot GEMBA-DA prompting, the model returning a 0-100 score per segment; within-language correlation averaged over nine directions (six English to Indic, three into English) of the 12,730-segment challenge test; each model over the segments it answers in the required format). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1GPT-5.50.62
2Sarvam M0.44
3aya-expanse-32B0.4
4Llama 3.2 3B Instruct0.22

Interactive version: theaggregate.ai/benchmark?slug=indicqe-ape-segment-level-quality-estimation · How It Works · Data refreshed daily, snapshot 2026-09-29.