Characterize the effect of \(\mathcal I_{\mathrm{opt}}\) on a principled exploration policy
Investigate the effect of applying the uncertainty-aware greedification operator \(\mathcal I_{\mathrm{opt}}\) to an exploration statistic such as an epistemic-uncertainty upper confidence bound when training a separate exploration policy \(\pi_e\).
References
We leave investigating the effect of such operator on the exploration policy to future work. However, we can predict that the operator has the same underlying effect: actions with certain high UCB will be favored, actions with certain low UCB will be unfavored, and actions with uncertain UCB will not be strongly modified in the exploration policy.
— Towards Optimal Policy Improvement
(2610.01566 - Oren et al., 1 Oct 2026) in Appendix, Section “How to utilize \(\mathcal I_{\mathrm{opt}}\) for principled exploration”