Math-Vision Diagrams - Edge F1: leaderboard
Metric: Edge F1 between Canny edge maps of the generated and reference diagram (0-1; mean over images on the 1,068 prompts whose output every one of the eleven models rendered; nine code LLMs write TikZ, SVG or Matplotlib code compiled to an image and two text-to-image models draw directly; provider default temperature, one response per prompt). Source: arxiv.org. Saturation forecast: Around 2032. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 0.17 |
| 2 | Claude Opus 4.6 | 0.16 |
| 3 | Kimi K2.5 | 0.14 |
| 4 | GPT-OSS-120B | 0.13 |
| 5 | Qwen 3.5 35B A3B | 0.11 |
| 6 | Llama 4 Maverick | 0.11 |
| 7 | GPT-5.4 | 0.08 |
Interactive version: theaggregate.ai/benchmark?slug=math-vision-diagrams-edge-f1 · How It Works · Data refreshed daily, snapshot 2026-09-29.