R2ABench (MetaGPT) - Faithfulness: leaderboard

Metric: Faithfulness score (1-5): judge rating (GPT-5.4-mini, calibrated against human review) of whether generated components, relations, technologies and assumptions are supported by the source evidence (L2), over the 68 R2ABench projects (17 educational and 51 GitHub repositories), generating a PlantUML architecture view from the structured requirements specification with MetaGPT adapted as an analyst, architect, reviewer and refiner agent pipeline under default settings and greedy decoding; rated over the L2-evaluable views, those that pass the L0 render-and-parse gate (the paper samples its judge calibration from the views that pass L0); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1MetaGPT + Claude Sonnet 4.62.23
2MetaGPT + GPT-52.17
3MetaGPT + DeepSeek V3.22.02
4MetaGPT + Qwen3-Coder 480B-A35B1.98

Interactive version: theaggregate.ai/benchmark?slug=r2abench-metagpt-faithfulness · How It Works · Data refreshed daily, snapshot 2026-10-07.