Mechanistic Effect of RL Post-Training on Pretrained Policies
Characterize how reinforcement learning with verifiable rewards changes the behavior of pretrained large language models for reasoning by determining whether RL primarily sharpens existing action preferences or discovers new, previously low-probability solution modes, and identify the conditions under which each behavior occurs.
References
As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model?
— Understanding Reasoning from Pretraining to Post-Training
(2607.16097 - Shen et al., 17 Jul 2026) in Abstract