Learning rulebooks from demonstrations and safety outcomes
Develop methods for learning rules, risk thresholds, and priority relations in risk-aware rulebooks from expert demonstrations or observed safety outcomes, so that the framework can adapt to different environments and operating conditions.
References
Several directions remain open for future work. First, the convergence analysis relies on conservative global Lipschitz constants and sampling assumptions. Developing weaker convergence conditions and adaptive local Lipschitz estimates could reduce conservatism. Approaches that avoid requiring a known Lipschitz constant altogether provide another promising direction . Second, the current certificates assume deterministic black-box evaluations of $q_i$. When $q_i$ is estimated from finite Monte Carlo samples or neural surrogates, statistical uncertainty and surrogate approximation error needs to be incorporated into the box bounds. Finally, learning different elements of rulebooks (e.g., rules, thresholds, priorities) from expert demonstrations or observed safety outcomes will enable the framework to adapt to different environments and operating conditions.