PM-LLM-Benchmark: leaderboard
Process mining benchmark evaluating LLMs on text-only tasks: contextual understanding of process data, conformance checking and anomaly detection, generating and modifying declarative and procedural process models, process querying, hypothesis generation, identifying unfairness, and process optimization. The diagram-interpretation category is defined upstream but is not part of the scored leaderboard, which reads event logs and process models as text.
Metric: Score. Source: github.com. Status: years away from saturation. 171 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (xHigh) | 40.2 |
| 2 | GPT-5.6 Sol (High) | 39.8 |
| 3 | Kimi K3 | 39.2 |
| 4 | GPT-5.6 Terra (High) | 38.4 |
| 5 | Kimi K2.6 | 37.8 |
| 6 | GPT-5.4 (xHigh) | 37.8 |
| 7 | GPT-5.5 (High) | 37.7 |
| 8 | DeepSeek V4 Flash (0731) | 37.7 |
| 9 | GPT-5.5 Pro | 37.7 |
| 10 | GPT-5.3 Codex (xHigh) | 37.3 |
| 11 | GPT-5.6 Luna (High) | 37.2 |
| 12 | GPT-5.4 (High) | 37 |
| 13 | GPT-5.5 (xHigh) | 36.9 |
| 14 | GPT-5 Pro | 36.9 |
| 15 | GPT-5.2 (xHigh) | 36.8 |
Interactive version: theaggregate.ai/benchmark?slug=pm-llm-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-05.