PAWBench (VLM Prediction) - Calibration: leaderboard

Metric: Conditional total-variation distance for vision-language models predicting outcomes as text without video generation, times 100, on a 0 to 100 scale (0-100); lower is better; PAWBench measures whether a system reproduces the reference distribution over physically possible outcomes across 50 rollouts on 25 scenarios per physical track, averaged over scenes that pass the readability gate. Source: arxiv.org. Saturation forecast: Around August 2028. 5 models tracked.

Top models

#ModelScore
1GLM-5V Turbo34.8
2Kimi K2.638.2
3Gemini 3.5 Flash39.4
4Qwen 3.5 Plus40.9
5GPT-5.542.3

Interactive version: theaggregate.ai/benchmark?slug=pawbench-vlm-prediction-calibration · How It Works · Data refreshed daily, snapshot 2026-09-26.