Assess MedAgent-R1 behavior at larger model scales
Evaluate whether the confident-hallucination phenomenon and the effectiveness of faithfulness-aware reinforcement learning persist for MedAgent-R1 backbones at scales substantially larger than 14B, particularly around 70B parameters or more.
References
Our primary experiments use a 7B backbone; a 14B replication (Appendix~\ref{app:scaling}) confirms that the phenomenon persists and the method transfers, but behavior at larger scales (70B+) remains untested.
— MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning
(2608.30676 - Chen et al., 31 Aug 2026) in Section ‘Limitations’