Math-Vision Diagrams - DISTS: leaderboard
Metric: DISTS perceptual distance between the generated and reference diagram (0-1, lower is better; mean over images on the 1,068 prompts whose output every one of the eleven models rendered; nine code LLMs write TikZ, SVG or Matplotlib code compiled to an image and two text-to-image models draw directly; provider default temperature, one response per prompt). Source: arxiv.org. Saturation forecast: Around May 2028. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 0.25 |
| 2 | Gemini 3.1 Pro (Preview) | 0.25 |
| 3 | GPT-OSS-120B | 0.26 |
| 4 | Kimi K2.5 | 0.26 |
| 5 | Llama 4 Maverick | 0.3 |
| 6 | Qwen 3.5 35B A3B | 0.32 |
| 7 | GPT-5.4 | 0.37 |
Interactive version: theaggregate.ai/benchmark?slug=math-vision-diagrams-dists · How It Works · Data refreshed daily, snapshot 2026-09-29.