FAM-Bench - Comparative Ranking (Baseline): leaderboard
Metric: Top-1 accuracy (%) on Task 2: rank four candidate dishes for a health condition and match the gold most suitable dish, Baseline prompting mode (the model answers directly), 1,000 comparative queries over four candidate dishes with recipe text and image, 13 health conditions, nutrition-expert verified, greedy decoding with JSON output; chance 25; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 32.7 |
| 2 | GPT-5.4 | 32.1 |
| 3 | Qwen 3 VL 8B Instruct | 31.3 |
| 4 | Claude Sonnet 4.6 | 30.9 |
| 5 | Gemma 3 12B (IT) | 30.6 |
Interactive version: theaggregate.ai/benchmark?slug=fam-bench-comparative-ranking-baseline · How It Works · Data refreshed daily, snapshot 2026-10-07.