R2ABench (Mini-SWE-Agent) - Faithfulness: leaderboard
Metric: Faithfulness score (1-5): judge rating (GPT-5.4-mini, calibrated against human review) of whether generated components, relations, technologies and assumptions are supported by the source evidence (L2), over the 68 R2ABench projects (17 educational and 51 GitHub repositories), generating a PlantUML architecture view from the structured requirements specification with the Mini-SWE-Agent framework under default settings and greedy decoding; rated over the L2-evaluable views, those that pass the L0 render-and-parse gate (the paper samples its judge calibration from the views that pass L0); higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 2.51 |
| 2 | Claude Sonnet 4.6 | 2.21 |
| 3 | DeepSeek V3.2 | 2.08 |
| 4 | Qwen 3 Coder 480B A35B | 2.06 |
Interactive version: theaggregate.ai/benchmark?slug=r2abench-mini-swe-agent-faithfulness · How It Works · Data refreshed daily, snapshot 2026-10-07.