TraceEval - Python: leaderboard

Metric: F1 (%) on the 769 Python test programs, edge-level F1 of the predicted call graph against execution-verified edges, zero-shot at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.679.1
2Claude Sonnet 4.671.3
3GPT-5.466.6
4DeepSeek V3.263.1
5GPT-5.4 Mini59.7
6Qwen 2.5 Coder 32B Instruct57.1
7Llama 3.3 70B Instruct53.3
8Gemini 3.1 Pro (Preview)49.5
9Qwen 2.5 Coder 7B Instruct27.6
10Qwen2.5-Coder-1.5B-Instruct16.1

Interactive version: theaggregate.ai/benchmark?slug=traceeval-python · How It Works · Data refreshed daily, snapshot 2026-10-07.