Reliable detection of reward hacking without success-rate evaluation
Develop reliable, scalable methods to detect over-optimization (reward hacking) during reinforcement learning from feedback of large language model agents trained with process reward models, without relying on task success-rate evaluation to assess true performance.
References
An open question remains how to reliably detect over-optimization without evaluating the success rate (which is difficult to scale).
— Process Reward Models for LLM Agents: Practical Framework and Directions
(2502.10325 - Choudhury, 14 Feb 2025) in Section 2.3 (Experiments), paragraph "Question: Can we measure and mitigate reward hacking?" adjacent to Figure "Process Reward Hacking"
Mitigating reward hacking in a way that scales with model capability is therefore an open and pressing problem.
— Debate Training Reduces Reward Hacking in RLAIF
(2608.17776 - Kenton et al., 18 Aug 2026) in Section 1, Introduction
An open problem is how to certify that evaluator fidelity survives policy optimization when success on RewardBench-style probes does not guarantee robustness to newly optimized outputs.
— Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
(2608.31075 - Yang et al., 31 Aug 2026) in Section 7.1, paragraph “Verifier robustness under optimization pressure”