Develop limited-feedback and non-stationary extensions

Develop extensions of the decentralized online continuous DR-submodular maximization framework to limited-feedback settings, including zeroth-order and bandit feedback, and characterize its performance in non-stationary environments through dynamic and adaptive regret.

Background

The paper’s application guarantees are established only for static regret under a natural first-order feedback model, in which each agent obtains a bounded first-order oracle response through an application-specific wrapper. The authors explicitly leave unresolved the adaptation of the framework to weaker feedback regimes, such as zeroth-order and bandit feedback, as well as to changing environments evaluated by dynamic or adaptive regret.

References

Separately, for the applications to DR-submodular classes, this work only discusses static regret under the natural first-order feedback setting. Extensions to limited feedback setting, including zeroth-order and bandit feedback, and the investigation of non-stationary environment like dynamic and adaptive regret, remain open.