PhotoBench (Multimodal Embedding): leaderboard
Metric: Recall@10 (%) on the Chinese queries, averaged over the three albums, where the model embeds the query text and the photos directly and ranks photos by similarity; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | RzenEmbed-v2-7B | 58 | |
| 2 | QQMM-embed-v2 | 57.8 | |
| 3 | Ops-MM-Embed-7B | 56.6 | |
| 4 | Qwen3-VL-Embed-8B | 53 | |
| 5 | VLM2Vec (PhotoBench checkpoint unspecified) | 52.4 | |
| 6 | Qwen3-VL-Embedding-2B | 50.6 | |
| 7 | B3-Qwen2-7B | 49.9 | |
| 8 | siglip2-giant-opt-patch16-256 | 47.6 | |
| 9 | siglip2-base-patch16-224 | 40 | |
| 10 | clip-ViT-B-32-multilingual-v1 | 6.1 |
Interactive version: theaggregate.ai/benchmark?slug=photobench-multimodal-embedding · How It Works · Data refreshed daily, snapshot 2026-10-11.