Reliability of the late-noise docking improvement

Determine whether the reported change in backward Bellman-gradient performance under the late-noise active-sensing docking condition is reliable and reproducible.

Background

The active-sensing docking task evaluates backward Bellman-gradient (BG) policy construction under regular and late-noise observation conditions. Although backward BG improves the reported docking cost and success rate relative to DPO in both conditions, the paper states that the late-noise change is unresolved. Thus, the repeatability or statistical reliability of the late-noise performance change remains an explicit unresolved issue.

References

The regular cost gain repeats on fresh episodes; the late-noise change is unresolved.

— Amortized Feedback Planning: Turning Model-Based Rollouts into Executable Policies  (2609.35012 - Huh, 28 Sep 2026) in Section 5.2, subsection “Active-sensing docking”

Under late noise, BG has lower cost with both maps, but the change in $\Delta$ is $+0.006442$ [$-0.007282,+0.020166$]; its direction is unresolved.

— Amortized Feedback Planning: Turning Model-Based Rollouts into Executable Policies  (2609.35012 - Huh, 28 Sep 2026) in Section 5.4, subsection “Matched policy-storage comparison”