TRAIL SWE — leaderboard
TRAIL (Trace Reasoning and Agentic Issue Localization) SWE subset: evaluates LLMs on trace-based debugging and issue localization in software engineering contexts.
Metric: Joint Accuracy. Source: huggingface.co. Status: years away from saturation. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro (Preview 05-06) | 5 |
| 2 | GPT-4.1 | 0 |
| 3 | Llama 4 Scout Instruct | 0 |
| 4 | Gemini 2.5 Flash (Preview 04-17) | 0 |
| 5 | Llama 4 Maverick Instruct | 0 |
Interactive version: theaggregate.ai/benchmark?slug=trail-swe · How the rankings work · Data refreshed daily, snapshot 2026-07-22.