Exact validity of the SOC gradient surrogate for trained controllers

Determine whether the stochastic optimal control identity nabla_{X_0}V(u,X_0;g)=-u_0(X_0) holds for finitely trained neural controllers used in Wasserstein and Sinkhorn Distributionally Robust Schrödinger Bridge adversarial updates, and whether it remains valid across other tasks and input distributions.

Background

The Wasserstein and Sinkhorn adversarial updates replace the gradient of the fixed-terminal-cost value function with the negative initial control, using an identity that is exact only for an optimal controller of the corresponding fixed-terminal-cost stochastic optimal control problem. In the implemented algorithm, however, the neural controller is obtained after finitely many optimization steps and is generally not known to be an exact SOC optimum.

The paper reports favorable empirical alignment between the computed cost gradient and the controller on one two-dimensional Gaussian experiment, but explicitly limits that evidence and does not establish exact optimality or transferability to other settings.

References

Thus, both direction and magnitude approximately satisfy the SOC relation on the evaluated samples, supporting its use in the adversarial update for this experiment. This is an empirical consistency check, not a proof of exact SOC optimality.

— Distributionally Robust Schrödinger Bridge  (2610.02043 - Sul et al., 1 Oct 2026) in Section 4, subsection “Additional Experimental Analysis,” paragraph “SOC optimality condition”; Appendix B, Section “Stochastic Optimal Control Theory”; Appendix H, subsection “Empirical Check of the SOC Optimality Condition”