AspectSim (Summarize-Then-Embed): leaderboard

Metric: Spearman correlation (-1 to 1) between the embedding similarity of the evidence the LLM extracts from each document and the GPT-4o aspect-similarity labels, mean over nine embedding models; the LLM writes an aspect-conditioned summary. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct0.57
2Phi-40.57
3Qwen 2.5 14B Instruct0.56
4Qwen 2.5 32B Instruct0.56
5Qwen 3 8B0.56
6DeepSeek R1 Distill Qwen 32B0.55
7Llama 3.3 70B Instruct0.54
8Gemma 3 12B (IT)0.53
9Qwen 3 4B0.52
10Gemma 2 27B (IT)0.52
11DeepSeek R1 Distill Qwen 14B0.48
12Mistral 7B Instruct0.47
13Llama 3.2 3B Instruct0.4
14Phi-4 Mini Instruct0.36
15Gemma 3 1B (IT)0.06

Interactive version: theaggregate.ai/benchmark?slug=aspectsim-summarize-then-embed · How It Works · Data refreshed daily, snapshot 2026-09-26.