PCB-Bench - Routing Macro QA BERTScore — leaderboard
Metric: BERTScore (%). Source: digailab.github.io. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3.1 | 83.1 |
| 2 | GPT-4o | 82.67 |
| 3 | MythoMax-L2-13B | 82.62 |
| 4 | GPT-5 | 81.93 |
| 5 | Llama 4 Maverick | 81.86 |
| 6 | Claude Opus 4.1 | 81.64 |
| 7 | InternVL3-78B | 80.71 |
| 8 | Gemini 2.5 Pro | 80.68 |
| 9 | Qwen 2.5 7B Instruct | 73.05 |
Interactive version: theaggregate.ai/benchmark?slug=pcb-bench-routing-macro-qa-bertscore · How the rankings work · Data refreshed daily, snapshot 2026-07-22.