Polynomial-time implementation with dimension-independent regret

Determine whether the $O(\sqrt d)$ expected-regret rate for online inverse linear optimization can be attained by an implementation whose running time is polynomial in the dimension, horizon, and input length.

Background

The paper studies online inverse linear optimization with a fixed unknown linear utility over compact action sets contained in the Euclidean unit ball. Its multiscale matrix-weights algorithm achieves an expected regret bound of O(d)O(\sqrt d) uniformly over the time horizon, and a rational-oracle implementation is shown to terminate in every round.

However, the implementation may require (dT)O(d)(dT)^{O(d)} linear-optimization oracle calls, together with additional finite computation. The paper therefore leaves unresolved whether the same optimal dependence on the dimension can be achieved while making the entire running time polynomial in the dimension, horizon, and input length.

References

Whether the same rate is attainable with running time polynomial in the dimension, horizon, and input length remains open.

However, this algorithm is inefficient and it remains an open question to obtain an efficient algorith with $O(\sqrt{d})$ regret.

— Cogentic: Multi-Agent Orchestration for Automated Proof Discovery  (2609.40324 - Cai et al., 30 Sep 2026) in Section 3, “Efficient Online Inverse Linear Optimization”