Occupation-measure relative-error condition for population OL-BPTT convergence

Establish the occupation-measure relative-error condition required to guarantee convergence of population fixed-latent OL-BPTT policy iteration for constrained CRRA portfolios, rather than verifying it only through a benchmark-specific audit.

Background

For constrained CRRA portfolio choice, the population OL-BPTT policy-improvement operator replaces the exact HJB normalized factor-gradient field Ru with the fixed-latent OL-BPTT field R_{\mathrm{OL}}u. The resulting adjoint–HJB Hamiltonian-gradient discrepancy is explicitly represented by C(R_{\mathrm{OL}}u-Ru).

The global convergence theorem does not follow from local residual decay alone. It assumes an occupation-measure relative-error inequality comparing the integrated adjoint–HJB defect with the integrated policy-update magnitude, with a gain-dependent constant below a specified threshold. The paper reports an empirical audit of this condition on a particular one-factor benchmark and visited iterate–start combinations, but does not establish the condition uniformly over all iterates and starting states required by the theorem.

References

What remains unverified is the occupation-measure relative error condition imposed later for population OL-BPTT convergence.

Self-Consistent Adjoint Policy Iteration for Constrained Dynamic Portfolio Choice  (2608.17808 - Huh et al., 18 Aug 2026) in Section 3, subsection “Adjoint and HJB portfolio Hamiltonians”