Qwen3-Embedding-8B: benchmark results
Provider: Alibaba. Access: Open.
Unified ELO 1538 ± 20, rank #631 of 1607 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BTZSC - Emotion | 50.69 | Macro-F1 (%) | 100 |
| SkMTEB - Retrieval | 86.27 | Mean score (%) over the five retrieval datasets; higher is b | 96.3 |
| BTZSC - Intent | 58.93 | Macro-F1 (%) | 88.2 |
| BTZSC - Sentiment | 89.16 | Macro-F1 (%) | 88.2 |
| SkMTEB | 74.53 | Mean score (%) across the 31 Slovak MTEB datasets over seven | 85.2 |
| SkMTEB - Reranking | 87.04 | Mean score (%) over the three reranking datasets; higher is | 85.2 |
| SkMTEB - STS | 86.54 | Mean score (%) over the two semantic-textual-similarity data | 85.2 |
| SkMTEB - Classification | 65.94 | Mean accuracy (%) over the seven classification datasets; hi | 81.5 |
| BTZSC | 59.13 | Macro-F1 (%) | 73.5 |
| SkMTEB - Pair Classification | 66.65 | Mean score (%) over the three pair-classification datasets; | 70.4 |
| SkillRet | 63.64 | NDCG@10 (0-100; first-stage dense retrieval of the gold skil | 64.7 |
| SkMTEB - Bitext Mining | 94.43 | Mean F1 (%) over the six bitext-mining datasets; higher is b | 63 |
Interactive version: theaggregate.ai/model?slug=qwen3-embedding-8b · How It Works · Data refreshed daily, snapshot 2026-09-29.