RFEval - Code Generation: leaderboard
Metric: Contrast-conditional reasoning faithfulness (%): after a counterfactual flawed reasoning segment is injected into the model's own thinking, the share of pairs in which the continued reasoning stays consistent with that stance and the stance causally carries into the final answer, computed only over the model's own contrast pairs (injection opposes its baseline stance; well-formed outputs); stances extracted by an o3 judge; greedy decoding; Code Generation task (861 items from LiveCodeBench and DS-1000); a property of reasoning traces rather than task accuracy; higher is better. Source: arxiv.org. 12 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | DeepSeek-R1-Distill-Qwen-7B | 38.25 | #1400 |
| 2 | DeepSeek R1 Distill Qwen 32B | 29.02 | #640 |
| 3 | DeepSeek R1 Distill Llama 70B | 27.89 | #582 |
| 4 | DeepSeek R1 Distill Llama 8B | 26.48 | #1282 |
| 5 | GPT-OSS-20B | 26.44 | #499 |
| 6 | Qwen 3 32B (Thinking) | 24.66 | #424 (Qwen 3 32B) |
| 7 | GPT-OSS-120B | 22.01 | #330 |
| 8 | Qwen 3 8B (Thinking) | 21.15 | #667 (Qwen 3 8B) |
Interactive version: theaggregate.ai/benchmark?slug=rfeval-code-generation · How It Works · Data refreshed daily, snapshot 2026-10-11.