Papers
Topics
Authors
Recent
Search
2000 character limit reached

Rectified Schrödinger Bridge Matching for Few-Step Visual Navigation

Published 7 Apr 2026 in cs.RO and cs.AI | (2604.05673v1)

Abstract: Visual navigation is a core challenge in Embodied AI, requiring autonomous agents to translate high-dimensional sensory observations into continuous, long-horizon action trajectories. While generative policies based on diffusion models and Schrödinger Bridges (SB) effectively capture multimodal action distributions, they require dozens of integration steps due to high-variance stochastic transport, posing a critical barrier for real-time robotic control. We propose Rectified Schrödinger Bridge Matching (RSBM), a framework that exploits a shared velocity-field structure between standard Schrödinger Bridges (ε=1\varepsilon=1, maximum-entropy transport) and deterministic Optimal Transport (ε0\varepsilon\to 0, as in Conditional Flow Matching), controlled by a single entropic regularization parameter ε\varepsilon. We prove two key results: (1) the conditional velocity field's functional form is invariant across the entire ε\varepsilon-spectrum (Velocity Structure Invariance), enabling a single network to serve all regularization strengths; and (2) reducing ε\varepsilon linearly decreases the conditional velocity variance, enabling more stable coarse-step ODE integration. Anchored to a learned conditional prior that shortens transport distance, RSBM operates at an intermediate ε\varepsilon that balances multimodal coverage and path straightness. Empirically, while standard bridges require 10\geq 10 steps to converge, RSBM achieves over 94% cosine similarity and 92% success rate in merely 3 integration steps -- without distillation or multi-stage training -- substantially narrowing the gap between high-fidelity generative policies and the low-latency demands of Embodied AI.

Summary

  • The paper introduces Rectified Schrödinger Bridge Matching, which combines a learned conditional prior with entropic bridge rectification to generate eight-waypoint navigation trajectories in as few as three ODE steps without distillation.
  • The method reduces velocity-target variance linearly with the regularization parameter ε while preserving the velocity-field structure, improving coarse-step integration and achieving 1.90 MSE, 0.945 cosine similarity, and 92% success at five function evaluations.
  • RSBM cuts function evaluations 3.8× versus a full-budget bridge baseline, generalizes across five navigation datasets, and runs at about 50 ms per decision cycle on a Jetson Orin, although very small ε values can reduce multimodal diversity.

Overview

"Rectified Schrödinger Bridge Matching for Few-Step Visual Navigation" (2604.05673) addresses a central deployment bottleneck for generative navigation policies: diffusion- and Schrödinger Bridge–based policies capture multimodal action distributions but require many integration steps because their Brownian transport paths are high-variance and curved. The authors propose Rectified Schrödinger Bridge Matching (RSBM), a single-stage framework that introduces an entropic regularization parameter ε(0,1]\varepsilon \in (0,1] into the bridge transition kernel, interpolating between maximum-entropy Schrödinger Bridges (ε=1\varepsilon = 1) and deterministic optimal-transport flow matching (ε0\varepsilon \to 0). The framework is anchored to a learned conditional prior that shortens the effective transport distance, enabling high-fidelity trajectory generation in as few as 3 ODE steps without distillation or multi-stage training.

Method

RSBM formulates visual navigation as conditional generative modeling: a dual-stream EfficientNet-B0/Transformer encoder maps streaming RGB observations and a goal image into a context vector cR256\mathbf{c} \in \mathbb{R}^{256}; a variational prior network gψg_\psi produces a coarse action initialization aT\mathbf{a}_T; and a conditional U-Net 1D velocity network vθ\mathbf{v}_\theta with FiLM conditioning refines aT\mathbf{a}_T into an 8-waypoint trajectory via a probability-flow ODE.

The core construction is the ε\varepsilon-rectified conditional bridge kernel. The mean follows the standard Brownian-bridge schedule μt=staT+(1st)a0\boldsymbol{\mu}_t = s_t \mathbf{a}_T + (1-s_t)\mathbf{a}_0 with ε=1\varepsilon = 10, while the variance is scaled by ε=1\varepsilon = 11: ε=1\varepsilon = 12. This preserves exact boundary pinning at both endpoints for any ε=1\varepsilon = 13. Setting ε=1\varepsilon = 14 recovers the standard Schrödinger Bridge; as ε=1\varepsilon = 15 the kernel collapses to the deterministic displacement interpolant of Monge–Kantorovich optimal transport. The paper grounds this parameterization in entropic optimal transport: identifying ε=1\varepsilon = 16 with the ratio of entropic regularization strengths yields a closed-form KL divergence between rectified and standard bridges of ε=1\varepsilon = 17, which is 1.55 nats at the default ε=1\varepsilon = 18 with ε=1\varepsilon = 19.

Training uses a simulation-free conditional flow matching loss on ε0\varepsilon \to 00-prediction targets, integrated at inference with a second-order Heun solver over a Karras schedule, giving NFE ε0\varepsilon \to 01 per ε0\varepsilon \to 02 steps.

Theoretical Results

Two formal results support the design.

Velocity Structure Invariance (Theorem 1): the logarithmic derivative ε0\varepsilon \to 03 is independent of ε0\varepsilon \to 04, since the ε0\varepsilon \to 05 factors cancel exactly. Consequently, the functional form of the conditional velocity field is invariant across the entire ε0\varepsilon \to 06-spectrum, so a single velocity-network parameterization serves all regularization strengths. In practice, ε0\varepsilon \to 07 acts only as a spatial support constrictor—concentrating training samples near the deterministic interpolant—rather than changing the learning target's structure.

Velocity Variance Reduction (Proposition 1): the conditional variance of the target velocity satisfies ε0\varepsilon \to 08, so reducing ε0\varepsilon \to 09 linearly reduces stochastic variation of the training target. The paper connects this to sampling error through a standard decomposition into approximation error and discretization error, arguing that lower target variance improves regression quality and yields smoother ODE right-hand sides amenable to coarse-step solvers. Notably, this connection is stated as consistent-with rather than derived-as: the paper does not prove a tight quantitative bound linking cR256\mathbf{c} \in \mathbb{R}^{256}0 to few-step accuracy, and it acknowledges that very small cR256\mathbf{c} \in \mathbb{R}^{256}1 causes over-regularization and degraded multimodal diversity.

The appendix also provides a signal-to-noise analysis motivating cR256\mathbf{c} \in \mathbb{R}^{256}2-prediction: unlike cR256\mathbf{c} \in \mathbb{R}^{256}3-prediction (which degrades near cR256\mathbf{c} \in \mathbb{R}^{256}4) or cR256\mathbf{c} \in \mathbb{R}^{256}5-prediction (which degrades near cR256\mathbf{c} \in \mathbb{R}^{256}6), the SNR of cR256\mathbf{c} \in \mathbb{R}^{256}7-prediction is well-behaved across the full interval, and it directly minimizes ODE integration error since the solver accumulates velocity predictions.

Empirical Results

Experiments span five public datasets (HuRoN, Recon, SACSoN, SCAND, GoStanford, ~60k trajectories total under ViNT/NoMaD splits), two simulation environments (Gazebo Custom Indoor, CitySim outdoor), and real-robot trials on a quadruped with a Jetson Orin. All methods are evaluated both at default step budgets and at cR256\mathbf{c} \in \mathbb{R}^{256}8 via zero-shot test-time step reduction, uniformly applied without retraining.

Method cR256\mathbf{c} \in \mathbb{R}^{256}9 NFE MSE↓ CosSim↑ Suc.%↑
DDPM 50 50 3.80 0.820 64
FM 10 10 2.80 0.910 82
NaviBridger 10 19 1.82 0.942 88
NaviBridger 3 5 12.00 0.710 35
RSBM 3 5 1.90 0.945 92
RSBM 10 19 1.72 0.949 93

(Custom Indoor; success/collision metrics from the paper's main comparison.)

Several results stand out. Few-step superiority: RSBM at gψg_\psi0 (NFE=5) matches or exceeds NaviBridger at its full budget of gψg_\psi1 (NFE=19)—a gψg_\psi2 reduction in function evaluations—with a 92% vs. 88% success rate and +4% over NaviBridger, while achieving gψg_\psi3 lower MSE than NaviBridger at gψg_\psi4. Under zero-shot step reduction, baselines degrade sharply: NaviBridger's CosSim falls from 0.942 to 0.710 and DDPM's to 0.320, whereas RSBM saturates early (MSE 1.90 → 1.72 from gψg_\psi5 to gψg_\psi6). Cross-dataset generalization: averaged over five real-world datasets at gψg_\psi7, RSBM attains MSE 1.19 / CosSim 0.934 versus NaviBridger(gψg_\psi8=10)'s 1.28 / 0.929, with the largest gains on GoStanford and SACSoN—long-range outdoor and dynamic-obstacle domains where path curvature amplifies truncation error. Prediction-target ablation: gψg_\psi9-prediction achieves 35.6% lower MSE than aT\mathbf{a}_T0-prediction and 45.7% lower than aT\mathbf{a}_T1-prediction at aT\mathbf{a}_T2, with all three converging by aT\mathbf{a}_T3, confirming the advantage is concentrated in the few-step regime. Real-robot deployment: RSBM runs at ~50 ms per decision cycle on a Jetson Orin, meeting a 4 Hz control rate, while DDPM (~350 ms) fails due to control-loop lag.

A four-way ablation disentangles the learned prior from bridge rectification: the prior alone yields MSE 5.8 (45% success); adding aT\mathbf{a}_T4-rectification lowers MSE to 1.9 (aT\mathbf{a}_T5); from Gaussian noise, RSBM still achieves MSE 4.2 versus 12.0 for standard SB (aT\mathbf{a}_T6), isolating rectification's contribution independent of prior quality. The two components are multiplicative rather than redundant. A solver ablation further shows Euler at aT\mathbf{a}_T7 already surpasses NaviBridger(aT\mathbf{a}_T8=10), indicating the advantage stems from the rectified bridge geometry rather than solver choice.

Limitations and Open Questions

The paper concedes several constraints. Simulation experiments evaluate closed-loop navigation, but real-world dataset results follow the open-loop offline protocol of prior work; the preliminary real-robot validation covers only a small number of indoor scenes without a standardized benchmark or dynamic obstacles. The learned prior limits zero-shot transfer to new environments. On the theoretical side, the variance-reduction result motivates but does not formally bound few-step discretization error as a function of aT\mathbf{a}_T9; the empirical tradeoff also shows that overly small vθ\mathbf{v}_\theta0 sacrifices multimodal coverage, and the choice vθ\mathbf{v}_\theta1 was tuned on one environment's validation set and held fixed elsewhere—an assumption about transferability of the operating point that is not independently validated per domain. Whether the intermediate-vθ\mathbf{v}_\theta2 sweet spot shifts across architectures, trajectory horizons, or out-of-distribution goals remains open.

Conclusion

RSBM unifies Schrödinger Bridges and flow matching through a single entropic parameter vθ\mathbf{v}_\theta3, proving that the conditional velocity field's structure is invariant across the family while its variance scales linearly with vθ\mathbf{v}_\theta4. Combined with a learned conditional prior, this yields a single-stage generative policy that reaches 94.5% cosine similarity and 92% success rate in 3 ODE steps, matching full-budget bridge baselines with vθ\mathbf{v}_\theta5 fewer function evaluations and no distillation. The framework provides a direct latency–quality knob for embodied deployment, though its real-world evidence remains limited to open-loop benchmarks and small-scale indoor robot trials.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.