Extensions of data-driven Brownian reflection control
Investigate four unresolved extensions of data-driven Brownian reflection control: a two-sided control framework; high-dimensional reflected Brownian control with analytically tractable optimal-policy characterization, effective learning algorithms, and regret analysis; finite-time regret analysis when explicit control costs are included; and data-dependent updating or stopping rules that replace the deterministic doubling schedule used by the adaptive-updating and full-history adaptive-updating algorithms.
References
Several directions remain open for future research. First, for reflection control, it would be natural to extend the present framework to a two-sided control framework. We expect not only our algorithm design principle but also the Foster-Lyapunov method employed in our transient regret analysis to remain implementable. Second, high-dimensional (reflected) Brownian control problems are substantially more challenging: on one hand, characterizing optimal policy is analytically intractable; on the other hand, designing high-dimensional learning algorithms and conducting regret analysis are more difficult. Third, it would be interesting to incorporate explicit control cost into the model, with the aim of conducting finite-time regret analysis. Finally, more flexible adaptive mechanisms may also be considered. The AU and AU-FH algorithms in our paper use a deterministic doubling schedule, which has two useful analytical consequences: the number of policy updates up to time $T$ is only $O(\log T)$, and the learning regret on each interval is bounded by a constant, which are critical to yield the $O(\log T)$ regret bound. Beyond such deterministic schedules, one may consider data-dependent updating or stopping rules, where the updating is adaptive to the information generated by the data.