AspectSim (Sentence-Level Retrieval): leaderboard

Metric: Spearman correlation (-1 to 1) between the embedding similarity of the evidence the LLM extracts from each document and the GPT-4o aspect-similarity labels, mean over nine embedding models; the LLM retrieves the single most relevant sentence. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct0.59
2Qwen 2.5 32B Instruct0.59
3Qwen 2.5 72B Instruct0.59
4Qwen 3 8B0.58
5Phi-40.58
6Qwen 2.5 14B Instruct0.58
7Gemma 2 27B (IT)0.57
8Qwen 3 4B0.57
9DeepSeek R1 Distill Qwen 32B0.57
10DeepSeek R1 Distill Qwen 14B0.57
11Gemma 3 12B (IT)0.55
12Mistral 7B Instruct0.52
13Phi-4 Mini Instruct0.48
14Llama 3.2 3B Instruct0.46
15Llama 3.2 1B Instruct0.11

Interactive version: theaggregate.ai/benchmark?slug=aspectsim-sentence-level-retrieval · How It Works · Data refreshed daily, snapshot 2026-09-26.