Principled reward design for reinforcement learning in autonomous driving
Develop a principled methodology for designing reward functions for reinforcement learning-based autonomous driving that effectively guides learning in dynamic traffic while avoiding suboptimal rule-based heuristics and aligning with evaluation metrics.
References
Due to the flexibility in optimization, there are many ways to define the reward for driving, making reward design an important open problem .
— CaRL: Learning Scalable Planning Policies with Simple Rewards
(2504.17838 - Jaeger et al., 24 Apr 2025) in Appendix, Section 'Related work', subsubsection 'Rewards for Driving'
Traffic-signal compliance is the weakest behavior reported: the policy stops correctly at a growing share of the red lights it meets as the curriculum proceeds, but still runs about a third of them.
— CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving
(2608.14332 - Saleem et al., 14 Aug 2026) in Section 5, Conclusion