FAM-Bench - Comparative Ranking (CoT): leaderboard
Metric: Top-1 accuracy (%) on Task 2: rank four candidate dishes for a health condition and match the gold most suitable dish, CoT prompting mode (the model first writes short reasoning steps), 1,000 comparative queries over four candidate dishes with recipe text and image, 13 health conditions, nutrition-expert verified, greedy decoding with JSON output; chance 25; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 35.7 |
| 2 | Gemini 2.5 Pro | 34.6 |
| 3 | GPT-5.4 | 33.1 |
| 4 | Gemma 3 12B (IT) | 31.3 |
| 5 | Qwen 3 VL 8B Instruct | 29.3 |
Interactive version: theaggregate.ai/benchmark?slug=fam-bench-comparative-ranking-cot · How It Works · Data refreshed daily, snapshot 2026-10-07.