Assess MedAgent-R1 behavior at larger model scales

Evaluate whether the confident-hallucination phenomenon and the effectiveness of faithfulness-aware reinforcement learning persist for MedAgent-R1 backbones at scales substantially larger than 14B, particularly around 70B parameters or more.

Background

The paper experimentally evaluates MedAgent-R1 primarily with a 7B backbone and reports a 14B replication showing that outcome-only reinforcement learning still degrades faithfulness while faithfulness-aware reinforcement learning improves it. However, the authors explicitly state that behavior at larger scales, including 70B and above, has not been tested. Determining whether scaling changes the accuracy–faithfulness trade-off or the effectiveness of the faithfulness-gated reward remains unresolved.

References

Our primary experiments use a 7B backbone; a 14B replication (Appendix~\ref{app:scaling}) confirms that the phenomenon persists and the method transfers, but behavior at larger scales (70B+) remains untested.

MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning  (2608.30676 - Chen et al., 31 Aug 2026) in Section ‘Limitations’