Mechanistic Effect of RL Post-Training on Pretrained Policies
Characterize how reinforcement learning with verifiable rewards changes the behavior of pretrained large language models for reasoning by determining whether RL primarily sharpens existing action preferences or discovers new, previously low-probability solution modes, and identify the conditions under which each behavior occurs.
References
As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model?
While we have validated SignBalance across settings (math reasoning and search-agent) and across model scales (0.5B, 3B, and 7B), and we empirically observe that our method drives the policy toward more effective reasoning strategies, we have not yet quantitatively measured the resulting growth in reasoning ability, such as the policy's reasoning confidence and its uncertainty behavior across problem difficulty levels.