Papers
Topics
Authors
Recent
Search
2000 character limit reached

Self-Consistent Adjoint Policy Iteration for Constrained Dynamic Portfolio Choice

Published 18 Aug 2026 in math.OC, q-fin.CP, and q-fin.PM | (2608.17808v1)

Abstract: We develop simulation-based policy iteration for continuous-time portfolio choice with predictable returns and convex constraints. Each outer step re-evaluates a fixed-latent OL-BPTT adjoint after deployment and solves the constrained update. Shifted-adjoint cancellation controls the adjoint--HJB Hamiltonian-gradient discrepancy by the policy-improvement residual. For CRRA portfolios, exact HJB policy iteration identifies the optimal reduced value factor, while population OL-BPTT iteration converges globally under an occupation-measure relative-error condition. A theorem-matched audit yields a maximal 95% upper endpoint of 0.074 against the required 0.75 threshold. In a three-factor, fifty-asset design, current-policy re-evaluation outperforms matched pooled refinement under both evaluation laws.

Summary

  • The paper develops a simulation-based policy iteration that repeatedly re-evaluates adjoint fields under the deployed policy, solves constrained control updates, and uses shifted-adjoint cancellation to avoid requiring exact value-Hessian recovery.
  • The paper proves local bounds linking Hamiltonian-gradient and policy-update errors to the policy-improvement residual, while establishing global convergence for constrained CRRA portfolios under an occupation-measure relative-error condition.
  • The paper shows that policy re-evaluation can substantially improve accuracy in predictable-return settings, reducing RMSE from 5.99×10⁻² to 6.58×10⁻⁴ in one experiment, although gains depend on representation and evaluation coverage.

Overview and contribution

This paper develops a simulation-based policy iteration for continuous-time portfolio choice with predictable returns and convex trading constraints. The method builds on a companion one-shot construction—fixed-latent open-loop backpropagation through time (OL-BPTT) reverse differentiation of a deployed rollout, conditional reconstruction of Pontryagin adjoints at state–time restart points ("anchors"), and an explicit constrained local control solve—and closes that map into a self-consistent outer iteration: each step re-evaluates the adjoint field under the currently deployed (damped) policy, solves the constrained action update, and deploys the result. The central object is the three-operator hierarchy TkAT\mathcal T_k \to A \to T, separating finite-computation error (Tk\mathcal T_k), population-level adjoint-based improvement (AA), and exact HJB policy improvement (TT).

The paper's structural insight is shifted-adjoint cancellation: when the control enters the diffusion coefficient, the shifted martingale input ζu=ZuPuσ(t,Stu,ut)\zeta^u = \mathsf Z^u - P^u\sigma(t,S_t^u,u_t) yields an exact decomposition in which the second-adjoint discrepancy PuDss2VuP^u - D^2_{ss}V^u appears only multiplied by the diffusion displacement (σ(a)σ(u))(\sigma(a)-\sigma(u)) of the trial action. This term vanishes at the current control and is proportional to the updated-control displacement on a regular policy step, so the full adjoint–HJB Hamiltonian-gradient discrepancy is controlled by the policy-improvement residual without requiring PuDss2VuP^u \to D^2_{ss}V^u. For CRRA portfolios, homotheticity removes the second-adjoint blocks entirely: only the normalized factor-gradient field ROLuR_{\mathrm{OL}}^u must be estimated, after which the many-asset action is recovered by an explicit metric projection or quadratic program.

Adjoint–HJB consistency

The consistency layer proceeds in stages. A fixed-policy value-gradient identification proposition gives the exact closed-loop reference and an explicit, auditable defect for the fixed-latent OL-BPTT target, which drops the state derivative of the deployed feedback from the reverse-mode chain rule. Under a factorized quadratic HJB Hamiltonian with positive weight wuw^u, envelope cancellation ties this first-order defect to the policy-improvement residual: on interior regions or fixed regular active faces with tangential KKT stationarity,

Tk\mathcal T_k0

so fixed points of Tk\mathcal T_k1 have zero first-order OL-BPTT defect. The main theorem then upgrades this to the complete Hamiltonian-gradient discrepancy at the updated control: on a regular future tube, under parabolic regularity and moment assumptions, both Tk\mathcal T_k2 and Tk\mathcal T_k3 are bounded by constants times the future-tube residual Tk\mathcal T_k4. The paper is explicit that this is a local statement: it does not verify the small relative-error constant needed globally, makes no claim across switching surfaces or moving feasible fibers, and uniform-over-action convergence of the Hamiltonian gradient is generally false without Tk\mathcal T_k5.

A boundary-stability proposition quantifies how pre-projection-field error propagates to switching geometry: under a transversality condition, Hausdorff error in the switching boundary is bounded by Tk\mathcal T_k6, and margin-type conditions yield misclassification rates Tk\mathcal T_k7 from Tk\mathcal T_k8 field error. This motivates the interpolate-then-project output rule: averaging pre-projection fields and projecting once preserves active geometry, whereas projecting damped averages blurs switching boundaries.

Value improvement and global convergence

For strongly concave quadratic Hamiltonians over convex feasible sets, the paper derives a damped value-improvement certificate: the exact half-step-style update satisfies

Tk\mathcal T_k9

and an approximate version decomposes the Hamiltonian-gradient error into five auditable components (consistency, discretization, statistics, representation, continuation extension) plus a KKT gap, yielding a certificate gain AA0 selected from the a posteriori lower bound. Residual summability follows without contraction, but the paper concedes that summability alone does not identify a unique global policy.

Global convergence is proved for the constrained CRRA subclass. Exact HJB policy iteration constructs the optimal reduced value factor AA1 via monotone limits with policy-uniform AA2 bounds; weighted Lyapunov-barrier comparison gives uniqueness within the positive strong-solution class AA3; and reduced verification identifies the limit with the portfolio value. The population OL-BPTT iteration converges globally under an occupation-measure relative-error condition AA4 with AA5: damping enlarges the tolerated ratio (half-step protocol permits AA6; full steps require AA7) but does not remove the population adjoint–HJB discrepancy itself.

After regular active-face identification, the exact full-step operator is locally quadratic in policy distance—the mechanism being envelope cancellation of the evaluation operator's derivative on the identified face—giving AA8 for the value gap under a separate coercivity condition on the weighted occupation norm. A quadratically accurate population OL-BPTT operator inherits this rate; fixed damping AA9 leaves only linear convergence unless TT0.

Numerical results

The experiments isolate when re-evaluation matters. In the Merton negative control (policy-independent improvement operator), concentrating budget in one shot yields RMSE TT1 versus TT2 for current-policy iteration at one hundred assets—iteration is not a universal variance-reduction device. With predictable returns, re-evaluation under the recovered policy reduces RMSE from TT3 to TT4, while spending the same matched budget at the initial policy leaves the error near TT5: the correction is genuinely policy-dependent information.

Audits support the theory's predictions. On a deterministic scaffold of exact damped HJB iterates, the log–log fit of defect against residual has slope TT6 with correlation TT7 on theorem-covered cells. The theorem-matched occupation-measure audit crosses four population iterates, three starting times, five initial factor states, and three seeds (180 banks): the largest per-bank 95% upper confidence endpoint for TT8 is 0.073936, roughly one tenth of the required half-step threshold 0.75. The authors note this audits only the visited sequence and listed starts, not the iterate-uniform condition.

In the three-factor, fifty-asset constrained benchmark with heterogeneous box constraints, current-policy re-evaluation beats matched pooled initial-policy refinement under both on-policy and broad evaluation laws in all three seeds (e.g., on-policy RMSE TT9 vs. ζu=ZuPuσ(t,Stu,ut)\zeta^u = \mathsf Z^u - P^u\sigma(t,S_t^u,u_t)0, a 37.69% reduction, under the legacy design). Representation matters sharply: quadratic features return near their oracle-fit benchmark (no benefit from iteration), while dense tanh features achieve roughly two-thirds one-shot error reduction. Design-point coverage is decisive—zero broad anchors degrade core RMSE to approximately 0.244—and the pre-specified sweep selects a 50–50 on-policy/broad mixture, improving broad and tail errors by about 41% and 44% relative to a 1/6 mixture. Notably, pooled refinement remains more accurate in the tail-only metric in all three seeds, so the advantage of re-evaluation is not uniform over the enlarged tail domain. An independent common-random-number finite-difference audit also exposed and corrected a factor-of-two error in an initial HJB reference stencil, illustrating the value of cross-validation against non-oracle references.

Limitations and open questions

The paper states its scope plainly. Global convergence holds only for the constrained CRRA population subclass; the generic consistency theorem is local to regular future tubes and does not verify the iterate-uniform relative-error constant. The audit covers a single benchmark sequence. Local quadratic rates assume prior finite identification of the active face, which is not proved; coercivity of the weighted occupation norm is assumed separately and may be strong on unbounded diffusions. No dimension-free factor-state complexity bound is claimed, and a fixed biased sampled implementation need not converge exactly to ζu=ZuPuσ(t,Stu,ut)\zeta^u = \mathsf Z^u - P^u\sigma(t,S_t^u,u_t)1—it is covered only by the a posteriori certificate. Open questions include whether the pointwise route to the relative-error condition can be established without a localization or residual-comparability estimate, whether the rate extends across switching surfaces, and what joint schedules over iterations, samples, and mesh sizes guarantee convergence of the sampled algorithm.

Conclusion

The paper closes a fixed-latent adjoint-to-control map into a self-consistent policy iteration whose theoretical guarantees rest on shifted-adjoint cancellation rather than value-Hessian identification. Global convergence for constrained CRRA portfolios follows from an occupation-measure relative-error condition that the numerical audit satisfies with substantial empirical margin, and the matched-budget experiments establish precisely when re-evaluation under the deployed policy—rather than pooled estimation under the initial policy—is the productive use of simulation budget.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.