Learn Potential-Based Shaping Useful Throughout Training
Develop a procedure to learn a potential function Φ(s) for potential-based reward shaping such that, when added to a recovered reward r to form r′(s,a)=r(s,a)+Φ(s′)−Φ(s), the shaped reward accelerates policy training from scratch throughout the entire training process, including early stages with weak initial policies.
References
Thus, in practice, we are left with an open question. Challenge 3: In practice, how do we learn a potential-based shaping term that is useful throughout the course of training from scratch?
— EvIL: Evolution Strategies for Generalisable Imitation Learning
(2406.11905 - Sapora et al., 2024) in Section “Reward-Centric Challenges of Efficient IRL”, Challenge 3
Whether Qwen-2.5:14b's shaping accelerates or decelerates learning in practice is an empirical question left to future work; the pilot in Section~\ref{sec:pipeline-validation} runs without shaping for CPU tractability.
— Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
(2608.18008 - Hounwanou et al., 18 Aug 2026) in Section 6, Discussion, subsection “What Proposition 1 does and does not guarantee”; reiterated in Section 8, Conclusion