Determine whether nonlinear KL-based distance shaping improves learning beyond direct distance guidance
Determine whether the normalized nonlinear distance penalty induced by the exact two-outcome KL proxy improves reinforcement-learning performance beyond direct successor-state distance guidance when observations, coefficient-selection opportunities, and training budgets are matched, and determine whether any performance difference depends on a constant cost incurred until episode termination.
References
The next empirical question is whether a nonlinear distance penalty improves learning beyond direct distance guidance when observations, coefficient-selection opportunities, and training budgets are matched. The associated mechanism question is whether any difference depends on a constant cost paid until termination.
— IncentRL: The Trade-Off Between Preference Guidance and Task Performance
(2609.21525 - Wu et al., 18 Sep 2026) in Section 5.1, “A prospective test of local guidance and task completion” (Section \ref{sec:prospective})