Math-Vision Diagrams - Edge F1: leaderboard

Metric: Edge F1 between Canny edge maps of the generated and reference diagram (0-1; mean over images on the 1,068 prompts whose output every one of the eleven models rendered; nine code LLMs write TikZ, SVG or Matplotlib code compiled to an image and two text-to-image models draw directly; provider default temperature, one response per prompt). Source: arxiv.org. Saturation forecast: Around 2032. 11 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)0.17
2Claude Opus 4.60.16
3Kimi K2.50.14
4GPT-OSS-120B0.13
5Qwen 3.5 35B A3B0.11
6Llama 4 Maverick0.11
7GPT-5.40.08

Interactive version: theaggregate.ai/benchmark?slug=math-vision-diagrams-edge-f1 · How It Works · Data refreshed daily, snapshot 2026-09-29.