Characterize the effect of \(\mathcal I_{\mathrm{opt}}\) on a principled exploration policy

Investigate the effect of applying the uncertainty-aware greedification operator \(\mathcal I_{\mathrm{opt}}\) to an exploration statistic such as an epistemic-uncertainty upper confidence bound when training a separate exploration policy \(\pi_e\).

Background

The paper proposes using policy-improvement operators to train an exploration policy with respect to an exploration statistic rather than the ordinary value or advantage. A specific candidate is an upper-confidence-bound statistic combining estimated value with epistemic uncertainty.

Although the authors describe the expected qualitative behavior of Iopt\mathcal I_{\mathrm{opt}}—favoring actions with certain high upper-confidence bounds and modifying uncertain upper-confidence bounds less strongly—they explicitly leave the empirical and theoretical investigation of its effect on the resulting exploration policy unresolved.

References

We leave investigating the effect of such operator on the exploration policy to future work. However, we can predict that the operator has the same underlying effect: actions with certain high UCB will be favored, actions with certain low UCB will be unfavored, and actions with uncertain UCB will not be strongly modified in the exploration policy.

— Towards Optimal Policy Improvement  (2610.01566 - Oren et al., 1 Oct 2026) in Appendix, Section “How to utilize \(\mathcal I_{\mathrm{opt}}\) for principled exploration”