- The paper establishes an O(d log d) regret bound independent of the time horizon for uncorrupted online inverse linear optimization over M-convex action sets, improving prior exponential finite bounds.
- The proposed method combines M-convex exchange optimality with a center-of-gravity volume argument, while a simpler topological-sorting algorithm achieves O(d²) regret with lower computational demands.
- The paper extends the guarantee to adversarial corruption in C rounds, achieving O((C+1)d log d) regret through cycle detection and adaptive restarts, nearly matching the Ω(d) lower bound.
Problem setting and motivation
The paper studies online inverse linear optimization, also known as contextual recommendation, where a learner sequentially observes an agent's optimal actions xt∈argx∈Xtmax⟨w∗,x⟩ over time-varying feasible sets Xt, and must recommend actions that perform well under the agent's hidden objective vector w∗. Regret is the cumulative gap ∑t⟨w∗,xt−x^t⟩ between the agent's optimal values and those achieved under the learner's estimates w^t.
Prior work left a conspicuous gap: known upper bounds either depend on T (e.g., O(dlogT)) or are finite but exponentially large (exp(O(dlogd))), while a lower bound of Ω(d) holds (2602.01682). Whether a finite regret bound polynomial in d is achievable was open. This paper resolves it for a broad and natural class of action sets: M-convex sets, which subsume matroid bases (uniform matroids for Xt0-sets, graphic matroids for spanning trees) and their integer-lattice extensions such as bounded-count allocations Xt1.
The motivation for restricting to M-convex sets is structural rather than arbitrary: when feasible sets are curved (e.g., ellipsoids), a single observation pins down the direction of Xt2 via its normal cone, making the problem trivial; polyhedral combinatorial sets are the genuinely hard case, and M-convexity supplies enough exchange structure to make progress while covering many practical problems.
Warm-up: an Xt3-regret algorithm
The key structural tool is Murota's characterization of optimality on M-convex sets: Xt4 maximizes Xt5 over an M-convex Xt6 if and only if Xt7 for every feasible unit exchange Xt8 (2602.01682). The algorithm maintains a set Xt9 of index pairs w∗0 such that w∗1 is implied by past observations, chooses w∗2 consistent with all pairs in w∗3, and augments w∗4 whenever the observed action reveals a new ordering constraint.
Two facts drive the analysis. First, the directed graph w∗5 is always acyclic—otherwise w∗6 would satisfy a strict cyclic chain of inequalities—so a consistent w∗7 always exists via topological sort. Second, by the optimality characterization, if no new pair is added at round w∗8 (i.e., w∗9), then ∑t⟨w∗,xt−x^t⟩0 is exactly the unique maximizer under ∑t⟨w∗,xt−x^t⟩1 (uniqueness follows from distinct components plus the M-convex exchange property). Hence nonzero regret occurs only when a new pair is learned, which can happen at most ∑t⟨w∗,xt−x^t⟩2 times, yielding ∑t⟨w∗,xt−x^t⟩3 (2602.01682).
This already improves on the exponential finite bound of prior work, but the choice of ∑t⟨w∗,xt−x^t⟩4 is still arbitrary; refining that choice yields the main result.
The ∑t⟨w∗,xt−x^t⟩5 bound via a volume argument
The refined algorithm takes ∑t⟨w∗,xt−x^t⟩6 as the center of gravity of the order-constrained polytope
∑t⟨w∗,xt−x^t⟩7
with ties broken by arbitrarily small perturbation. Whenever a mistake occurs (∑t⟨w∗,xt−x^t⟩8), some newly added pair ∑t⟨w∗,xt−x^t⟩9 satisfies w^t0; since w^t1 lies outside the halfspace w^t2, Grünbaum's theorem implies w^t3 (2602.01682).
To bound the total number of mistakes, the paper uses the Freudenthal triangulation: w^t4 decomposes into w^t5 equal-volume order simplices, and because all constraints in w^t6 are consistent with the true ordering of w^t7's components, w^t8 contains at least one full order simplex. Thus at most w^t9 mistakes occur, giving:
Main result: under M-convex action sets and uncorrupted feedback, regret T0, independent of T1 (2602.01682).
A practical caveat concerns computation: exact center-of-gravity computation is #P-hard even for order simplices, so the algorithm requires randomized approximation of the center of gravity in polynomial time; the volume analysis tolerates this via approximate versions of Grünbaum's theorem. The simpler T2 variant needs only topological sorting, costing T3 per round.
Corruption-robust extension
The paper then relaxes the assumption that observed actions are optimal, allowing adversarial corruption in up to T4 rounds (the standard corruption model from contextual search/pricing), without assuming the learner knows T5. The regret definition uses actual actions; the alternative definition based on unobserved optimal actions differs by only constant factors, since T6 and T7 in the worst case.
The mechanism rests on two lemmas about the directed graph induced by feedback over any interval T8 (2602.01682):
- Cycle implies corruption: if the graph built from arc sets T9 over the interval contains a directed cycle, at least one action in the interval is suboptimal.
- Acyclicity plus no new arcs implies zero regret: if the interval graph is acyclic and no new arcs appear at round O(dlogT)0, then O(dlogT)1 regardless of earlier corruptions.
The algorithm runs the uncorrupted procedure and restarts (resetting O(dlogT)2) whenever a cycle appears. Each inter-restart epoch accumulates at most O(dlogT)3 regret by the same volume argument—the volume-decrease lemma does not require optimality of observed actions—and each cycle certifies at least one corrupted round, so there are at most O(dlogT)4 restarts. The resulting guarantee is O(dlogT)5, recovering the clean bound at O(dlogT)6 and requiring no prior knowledge of O(dlogT)7 (2602.01682). Cycle detection adds only O(dlogT)8 per-round overhead via DFS or topological sort.
Compared with Gupta et al.'s parallel-copies framework for contextual pricing, which also incurs multiplicative O(dlogT)9 degradation, this approach instead exploits M-convex structure for detection and uses restarts rather than maintaining multiple algorithm copies. One limitation the authors note explicitly: Sakaue et al.'s more flexible corruption model, measured by cumulative suboptimality rather than round counts, is not covered; extending the finite-regret approach to that framework remains open.
Lower bound
The paper adapts the exp(O(dlogd))0 lower-bound construction of Sakaue et al., which uses axis-aligned line-segment feasible sets, to the M-convex setting (2602.01682). The segment instances are not M-convex but are Mexp(O(dlogd))1-convex (box-type sets); since any Mexp(O(dlogd))2-convex set in exp(O(dlogd))3 embeds as an M-convex set in exp(O(dlogd))4, the lower bound carries over at dimension exp(O(dlogd))5. Consequently, the exp(O(dlogd))6 upper bound is tight up to a logarithmic factor, and the gap between the bounds is now only exp(O(dlogd))7.
Limitations and open questions
Several caveats qualify the results. The exp(O(dlogd))8 guarantee relies on randomized approximate center-of-gravity computation due to #P-hardness of the exact quantity. The corruption guarantee counts corrupted rounds rather than cumulative suboptimality, matching the weaker of the two models in the literature. The boundedness assumption (per-round regret exp(O(dlogd))9) and the distinct-components assumption on Ω(d)0 are standard but simplifying; the latter is handled via perturbation arguments justified by Lipschitz continuity of regret. Open questions identified by the authors include closing the Ω(d)1 gap, determining the tight dependence on Ω(d)2, achieving corruption-robust bounds under cumulative-suboptimality measures, and understanding what additional structure on Ω(d)3 and the action sets permits learning faster than the Ω(d)4 barrier—a question connected to preference-feedback settings, which two-action M-convex sets model as a special case.
Conclusion
This paper establishes that M-convexity of the action sets suffices for finite, dimension-polynomial regret in online inverse linear optimization: Ω(d)5 without corruptions and Ω(d)6 with adversarial corruptions detected adaptively, against an Ω(d)7 lower bound. The technical contribution is an overview of Murota's local-exchange optimality characterization with a Grünbaum-style volume argument over order simplices, extended to corruptions through acyclicity monitoring of feedback-induced directed graphs. The results partially close a question open since Gollapudi et al. and Sakaue et al., and position discrete convexity as a productive structural assumption for online inverse optimization.