Extensions of data-driven Brownian reflection control

Investigate four unresolved extensions of data-driven Brownian reflection control: a two-sided control framework; high-dimensional reflected Brownian control with analytically tractable optimal-policy characterization, effective learning algorithms, and regret analysis; finite-time regret analysis when explicit control costs are included; and data-dependent updating or stopping rules that replace the deterministic doubling schedule used by the adaptive-updating and full-history adaptive-updating algorithms.

Background

The paper develops learn-then-optimize, adaptive-updating, and full-history adaptive-updating algorithms for one-sided Brownian reflection control with unknown drift and volatility. Their regret analysis relies on the one-dimensional structure of the reflected Brownian motion, the explicitly characterized optimal reflection level, Foster–Lyapunov estimates, and a deterministic doubling schedule for policy updates.

The concluding section identifies unresolved extensions beyond this setting. These include allowing control on both sides of the state, addressing high-dimensional reflected Brownian control where optimal policies and learning guarantees are difficult to characterize, incorporating explicit control costs into finite-horizon regret, and replacing deterministic update times with rules driven by observed data.

References

Several directions remain open for future research. First, for reflection control, it would be natural to extend the present framework to a two-sided control framework. We expect not only our algorithm design principle but also the Foster-Lyapunov method employed in our transient regret analysis to remain implementable. Second, high-dimensional (reflected) Brownian control problems are substantially more challenging: on one hand, characterizing optimal policy is analytically intractable; on the other hand, designing high-dimensional learning algorithms and conducting regret analysis are more difficult. Third, it would be interesting to incorporate explicit control cost into the model, with the aim of conducting finite-time regret analysis. Finally, more flexible adaptive mechanisms may also be considered. The AU and AU-FH algorithms in our paper use a deterministic doubling schedule, which has two useful analytical consequences: the number of policy updates up to time $T$ is only $O(\log T)$, and the learning regret on each interval is bounded by a constant, which are critical to yield the $O(\log T)$ regret bound. Beyond such deterministic schedules, one may consider data-dependent updating or stopping rules, where the updating is adaptive to the information generated by the data.

— Data-Driven Brownian Reflection Control  (2609.03870 - Pang et al., 3 Sep 2026) in Section 8, Concluding Remarks