NutriBench-Kitchen: leaderboard
Metric: Average score (%; unweighted mean of the ten task-regime cells: five task families (ingredient entry, memory management, recipe query, long-term and short-term planning) in an easy and a hard regime; NutriBench-Kitchen: 1,500 manually verified questions on 160 cooking videos (145 HD-EPIC segments and 15 self-recorded videos with measured ingredient quantities); accuracy on multiple-choice questions and mean relative accuracy on numerical ones; zero-shot with the same chain-of-thought prompt for every model). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 61.1 |
| 2 | Qwen 3 VL 8B | 53.4 |
| 3 | Qwen 2.5 VL 7B Instruct | 50.8 |
| 4 | GPT-4o | 47.7 |
| 5 | GLM-4.1V-9B (Thinking) | 29.4 |
Interactive version: theaggregate.ai/benchmark?slug=nutribench-kitchen · How It Works · Data refreshed daily, snapshot 2026-09-26.