- The paper develops a simulation-based policy iteration that repeatedly re-evaluates adjoint fields under the deployed policy, solves constrained control updates, and uses shifted-adjoint cancellation to avoid requiring exact value-Hessian recovery.
- The paper proves local bounds linking Hamiltonian-gradient and policy-update errors to the policy-improvement residual, while establishing global convergence for constrained CRRA portfolios under an occupation-measure relative-error condition.
- The paper shows that policy re-evaluation can substantially improve accuracy in predictable-return settings, reducing RMSE from 5.99×10⁻² to 6.58×10⁻⁴ in one experiment, although gains depend on representation and evaluation coverage.
Overview and contribution
This paper develops a simulation-based policy iteration for continuous-time portfolio choice with predictable returns and convex trading constraints. The method builds on a companion one-shot construction—fixed-latent open-loop backpropagation through time (OL-BPTT) reverse differentiation of a deployed rollout, conditional reconstruction of Pontryagin adjoints at state–time restart points ("anchors"), and an explicit constrained local control solve—and closes that map into a self-consistent outer iteration: each step re-evaluates the adjoint field under the currently deployed (damped) policy, solves the constrained action update, and deploys the result. The central object is the three-operator hierarchy Tk→A→T, separating finite-computation error (Tk), population-level adjoint-based improvement (A), and exact HJB policy improvement (T).
The paper's structural insight is shifted-adjoint cancellation: when the control enters the diffusion coefficient, the shifted martingale input ζu=Zu−Puσ(t,Stu,ut) yields an exact decomposition in which the second-adjoint discrepancy Pu−Dss2Vu appears only multiplied by the diffusion displacement (σ(a)−σ(u)) of the trial action. This term vanishes at the current control and is proportional to the updated-control displacement on a regular policy step, so the full adjoint–HJB Hamiltonian-gradient discrepancy is controlled by the policy-improvement residual without requiring Pu→Dss2Vu. For CRRA portfolios, homotheticity removes the second-adjoint blocks entirely: only the normalized factor-gradient field ROLu must be estimated, after which the many-asset action is recovered by an explicit metric projection or quadratic program.
Adjoint–HJB consistency
The consistency layer proceeds in stages. A fixed-policy value-gradient identification proposition gives the exact closed-loop reference and an explicit, auditable defect for the fixed-latent OL-BPTT target, which drops the state derivative of the deployed feedback from the reverse-mode chain rule. Under a factorized quadratic HJB Hamiltonian with positive weight wu, envelope cancellation ties this first-order defect to the policy-improvement residual: on interior regions or fixed regular active faces with tangential KKT stationarity,
Tk0
so fixed points of Tk1 have zero first-order OL-BPTT defect. The main theorem then upgrades this to the complete Hamiltonian-gradient discrepancy at the updated control: on a regular future tube, under parabolic regularity and moment assumptions, both Tk2 and Tk3 are bounded by constants times the future-tube residual Tk4. The paper is explicit that this is a local statement: it does not verify the small relative-error constant needed globally, makes no claim across switching surfaces or moving feasible fibers, and uniform-over-action convergence of the Hamiltonian gradient is generally false without Tk5.
A boundary-stability proposition quantifies how pre-projection-field error propagates to switching geometry: under a transversality condition, Hausdorff error in the switching boundary is bounded by Tk6, and margin-type conditions yield misclassification rates Tk7 from Tk8 field error. This motivates the interpolate-then-project output rule: averaging pre-projection fields and projecting once preserves active geometry, whereas projecting damped averages blurs switching boundaries.
Value improvement and global convergence
For strongly concave quadratic Hamiltonians over convex feasible sets, the paper derives a damped value-improvement certificate: the exact half-step-style update satisfies
Tk9
and an approximate version decomposes the Hamiltonian-gradient error into five auditable components (consistency, discretization, statistics, representation, continuation extension) plus a KKT gap, yielding a certificate gain A0 selected from the a posteriori lower bound. Residual summability follows without contraction, but the paper concedes that summability alone does not identify a unique global policy.
Global convergence is proved for the constrained CRRA subclass. Exact HJB policy iteration constructs the optimal reduced value factor A1 via monotone limits with policy-uniform A2 bounds; weighted Lyapunov-barrier comparison gives uniqueness within the positive strong-solution class A3; and reduced verification identifies the limit with the portfolio value. The population OL-BPTT iteration converges globally under an occupation-measure relative-error condition A4 with A5: damping enlarges the tolerated ratio (half-step protocol permits A6; full steps require A7) but does not remove the population adjoint–HJB discrepancy itself.
After regular active-face identification, the exact full-step operator is locally quadratic in policy distance—the mechanism being envelope cancellation of the evaluation operator's derivative on the identified face—giving A8 for the value gap under a separate coercivity condition on the weighted occupation norm. A quadratically accurate population OL-BPTT operator inherits this rate; fixed damping A9 leaves only linear convergence unless T0.
Numerical results
The experiments isolate when re-evaluation matters. In the Merton negative control (policy-independent improvement operator), concentrating budget in one shot yields RMSE T1 versus T2 for current-policy iteration at one hundred assets—iteration is not a universal variance-reduction device. With predictable returns, re-evaluation under the recovered policy reduces RMSE from T3 to T4, while spending the same matched budget at the initial policy leaves the error near T5: the correction is genuinely policy-dependent information.
Audits support the theory's predictions. On a deterministic scaffold of exact damped HJB iterates, the log–log fit of defect against residual has slope T6 with correlation T7 on theorem-covered cells. The theorem-matched occupation-measure audit crosses four population iterates, three starting times, five initial factor states, and three seeds (180 banks): the largest per-bank 95% upper confidence endpoint for T8 is 0.073936, roughly one tenth of the required half-step threshold 0.75. The authors note this audits only the visited sequence and listed starts, not the iterate-uniform condition.
In the three-factor, fifty-asset constrained benchmark with heterogeneous box constraints, current-policy re-evaluation beats matched pooled initial-policy refinement under both on-policy and broad evaluation laws in all three seeds (e.g., on-policy RMSE T9 vs. ζu=Zu−Puσ(t,Stu,ut)0, a 37.69% reduction, under the legacy design). Representation matters sharply: quadratic features return near their oracle-fit benchmark (no benefit from iteration), while dense tanh features achieve roughly two-thirds one-shot error reduction. Design-point coverage is decisive—zero broad anchors degrade core RMSE to approximately 0.244—and the pre-specified sweep selects a 50–50 on-policy/broad mixture, improving broad and tail errors by about 41% and 44% relative to a 1/6 mixture. Notably, pooled refinement remains more accurate in the tail-only metric in all three seeds, so the advantage of re-evaluation is not uniform over the enlarged tail domain. An independent common-random-number finite-difference audit also exposed and corrected a factor-of-two error in an initial HJB reference stencil, illustrating the value of cross-validation against non-oracle references.
Limitations and open questions
The paper states its scope plainly. Global convergence holds only for the constrained CRRA population subclass; the generic consistency theorem is local to regular future tubes and does not verify the iterate-uniform relative-error constant. The audit covers a single benchmark sequence. Local quadratic rates assume prior finite identification of the active face, which is not proved; coercivity of the weighted occupation norm is assumed separately and may be strong on unbounded diffusions. No dimension-free factor-state complexity bound is claimed, and a fixed biased sampled implementation need not converge exactly to ζu=Zu−Puσ(t,Stu,ut)1—it is covered only by the a posteriori certificate. Open questions include whether the pointwise route to the relative-error condition can be established without a localization or residual-comparability estimate, whether the rate extends across switching surfaces, and what joint schedules over iterations, samples, and mesh sizes guarantee convergence of the sampled algorithm.
Conclusion
The paper closes a fixed-latent adjoint-to-control map into a self-consistent policy iteration whose theoretical guarantees rest on shifted-adjoint cancellation rather than value-Hessian identification. Global convergence for constrained CRRA portfolios follows from an occupation-measure relative-error condition that the numerical audit satisfies with substantial empirical margin, and the matched-budget experiments establish precisely when re-evaluation under the deployed policy—rather than pooled estimation under the initial policy—is the productive use of simulation budget.