---
title: Online Inverse Linear Optimization
url: https://www.emergentmind.com/topics/online-inverse-linear-optimization
type: topic
---

# Online Inverse Linear Optimization

Online Inverse Linear Optimization (OILO) studies the sequential recovery or emulation of an unknown linear optimization model from observed input–decision pairs. At each round, a decision maker receives a feasible region or optimization instance and selects an action that is optimal, approximately optimal, or behaviorally representative under an unknown objective or constraint parameter. The learner observes the instance and decision, updates its model, and may produce a recommendation before observing the current decision. OILO encompasses inverse optimization with online structured prediction, online learning of objective coefficients, contextual recommendation, online inverse optimal control, and sequential learning of feasible regions. Its central distinction from batch inverse optimization is that observations arrive incrementally and the learner must balance computational tractability, adaptation, identifiability, and decision regret.

## 1. Conceptual foundations and problem formulations

Inverse optimization reverses the usual direction of optimization. A forward model specifies an objective and feasible set and computes an optimizer; an inverse model receives a feasible region or optimization instance together with an observed or desired decision and infers parameters that make that decision optimal or approximately optimal.

For a maximization problem with feasible set $\mathcal F$, weight vector $w$, and value $c(w,S)$, a prescribed feasible solution $S$ is $\delta$-optimal when

$$
c(w,S)\geq c(w,S')+\delta,\qquad \forall S'\in\mathcal F,\;S'\neq S.
$$

The corresponding minimization inequality is reversed. The margin-based inverse problem changes the weights minimally:

$$
\min_{w'}\|w'-w\|_2^2
$$

subject to the prescribed solution being $\delta$-optimal. This formulation is central to online structured prediction because it corrects the model globally against every feasible competitor rather than only against the current prediction [1510.03130].

A general online inverse-optimization model considers a parameterized decision problem

$$
\min_x f(x,u,\theta)\quad\text{s.t.}\quad g(x,u,\theta)\leq 0,
$$

where $u$ is an observed signal, $\theta$ is an unknown preference or restriction parameter, and $S(u,\theta)$ is the optimal solution set. The learner observes sequential pairs $(u_t,y_t)$, where $y_t$ may be a noisy or infeasible observation of an optimal decision. The squared-distance loss is

$$
\ell(y,u,\theta)=\min_{x\in S(u,\theta)}\|y-x\|^2.
$$

An implicit online update is

$$
\theta_{t+1}
=
\arg\min_{\theta\in\Theta}
\left\{
\frac12\|\theta-\theta_t\|^2+\eta_t\ell(y_t,u_t,\theta)
\right\}.
$$

The update is implicit because the loss is evaluated at the new parameter. It does not require an explicit gradient or differentiability of the optimizer map [1810.01920].

A more specialized online inverse linear optimization setting assumes an unknown fixed objective $c^*$ and changing feasible sets $X_t$. The agent chooses

$$
x_t\in\arg\max_{x\in X_t}\langle c^*,x\rangle,
$$

while the learner predicts $\hat c_t$, recommends

$$
\hat x_t\in\arg\max_{x\in X_t}\langle\hat c_t,x\rangle,
$$

and then observes $x_t$. The principal decision-regret criterion is

$$
R_T^{c^*}
=
\sum_{t=1}^T
\langle c^*,x_t-\hat x_t\rangle.
$$

The surrogate regret is

$$
\widetilde R_T^{c^*}
=
\sum_{t=1}^T
\langle\hat c_t-c^*,\hat x_t-x_t\rangle,
$$

which decomposes into explanatory suboptimality and true-objective recommendation loss [2501.14349].

The learner generally cannot identify the numerical objective uniquely. Positive scaling leaves linear-program optimizers unchanged, and distinct objective vectors may lie in the same normal cone or induce identical decisions on all observed feasible sets. Consequently, OILO usually targets behavioral emulation and low decision regret rather than convergence of $\hat c_t$ to $c^*$ in norm.

## 2. Structured prediction and combinatorial inverse optimization

The earliest formulation considered here connects inverse combinatorial optimization to online structured prediction. For a training example $(x_t,y_t)$, each element or edge $e$ has feature vector $f_e$, and the model weight is

$$
w(e)=\theta_t^\top f_e.
$$

The current prediction is obtained from the forward combinatorial optimization problem,

$$
\hat y_t
=
\arg\max_{y\in\mathcal Y_t}
\theta_t^\top\phi(x_t,y),
$$

where $\phi(x_t,y)$ is the aggregate feature vector. The inverse update minimizes the change in $\theta_t$ subject to the labeled structure becoming $\delta_t$-optimal:

$$
\theta_{t+1}
=
\arg\min_{\theta'}
\|\theta'-\theta_t\|_2^2
$$

subject to

$$
\theta'^\top
\bigl(\phi(x_t,y_t)-\phi(x_t,y)\bigr)
\geq\delta_t,
\qquad
\forall y\neq y_t.
$$

The margin is typically chosen as the user-defined loss $\ell(y_t,\hat y_t)$. Unlike unconstrained inverse optimization, edge weights are coupled through $\theta'^\top f_e$ and cannot be perturbed independently. The method differs from one-best MIRA because it enforces the desired structure against all feasible structures using exact combinatorial optimality conditions [1510.03130].

For maximum-weight matroid basis, if $B$ is the prescribed basis and $C_B(f)$ is the unique circuit in $B\cup\{f\}$, $\delta$-optimality is equivalent to

$$
w'(e)-w'(f)\geq\delta,
\qquad
\forall f\notin B,\;\forall e\in C_B(f).
$$

This yields a polynomial-size convex quadratic program. A graphic matroid gives an inverse maximum-spanning-tree formulation.

For matroid intersection, a common basis is characterized through an exchange graph. A directed even cycle represents a multielement exchange, and its length measures the loss associated with the corresponding competing common basis. The prescribed basis is $\delta$-optimal if and only if the exchange graph contains no directed cycle of length at most $\delta$. Since explicitly imposing every cycle inequality can be exponential, distance-label variables provide a polynomial-size extended formulation.

The same cycle-elimination principle applies to perfect matching and minimum-cost maximum flow. In bipartite matching, alternating cycles correspond to directed cycles in an auxiliary graph; the absence of a $\delta$-augmenting alternating cycle is equivalent to $\delta$-optimality. For flows, a prescribed maximum flow is $\delta$-optimal if its residual graph has no cycle of weight less than $\delta$. Shortest paths and shortest-path trees can be handled through reductions to inverse arborescence or through distance variables enforcing a $\delta$ separation between tree and non-tree routes.

These formulations support applications including dependency parsing, word alignment, reviewer-to-paper assignment, and shortest-path prediction. Their common principle is that a global comparison over an exponential feasible family can be replaced by a compact combinatorial optimality characterization.

## 3. Online learning algorithms and computational mechanisms

Several algorithmic paradigms have been developed for OILO.

### Implicit KKT-based updates

The generalized online framework uses the proximal update above and replaces the lower-level optimality condition with KKT conditions. For a quadratic forward problem,

$$
\min_x \frac12x^\top Qx+c^\top x
\quad\text{s.t.}\quad
Ax\geq b_t,
$$

the online update introduces primal variables, KKT multipliers, and binary complementarity variables. The resulting problem is a mixed-integer second-order conic program. Linear programs admit analogous mixed-integer quadratic or conic reformulations. The method processes one observation at a time and retains only the current parameter estimate, although each update may itself be computationally difficult [1810.01920].

The theoretical analysis assumes a compact convex parameter domain, bounded feasible regions, strong convexity in the forward decision, Lipschitz parameter dependence, and approximate convexity of the optimizer map. Under these conditions, the regret relative to the best fixed batch parameter satisfies an $O(1/\sqrt T)$ average rate, and the method is statistically consistent under finite second moments. These guarantees do not automatically apply to arbitrary LPs because LP solution maps may be set-valued and discontinuous.

### Oracle-based online convex learning

For arbitrary feasible sets admitting a linear optimization oracle, the online loss

$$
f_t(c)
=
\max_{x\in X(p_t)}c^\top x-c^\top x_t
$$

is convex, with subgradient $\hat x_t-x_t$. This enables multiplicative weights and online gradient descent without KKT reformulations, and therefore covers mixed-integer and nonconvex feasible sets. Multiplicative weights is particularly effective for nonnegative normalized objectives, while online gradient descent accommodates general compact convex objective sets. Both yield $O(1/\sqrt T)$ average total-error guarantees under bounded-diameter assumptions [1810.01920].

The same loss is exactly a Fenchel–Young loss associated with the indicator function of the feasible set:

$$
L_X(c;x)
=
\max_{x'\in X}\langle c,x'\rangle-\langle c,x\rangle.
$$

This interpretation clarifies the distinction between explanatory suboptimality and recommendation quality. It also permits guarantees when the observed agent actions are not optimal: the learned objective can achieve explanatory loss close to that of the unknown objective, up to the online-learning term [2501.13648].

### Maximum optimality margin

Maximum optimality margin (MOM) is designed for contextual LPs and inverse LP when optimal decisions, but not objective coefficients, are observed. Under LP nondegeneracy, an observed solution reveals its optimal basis. The corresponding reduced-cost inequalities are imposed with a positive margin and slack variables:

$$
\hat r_{t,\mathcal N_t^*}
\geq
\mathbf 1-s_t,
\qquad
s_t\geq 0.
$$

For a contextual objective $\hat c_t=\Theta z_t$, the batch formulation is a convex quadratic program with Frobenius regularization and $\ell_1$ margin violations. In the online version, projected subgradient descent updates $\Theta_t$ after the observed optimal basis is revealed. Under bounded covariates and well-conditioned basis matrices, the method obtains an $O(T^{-1/2})$ average predicted-objective suboptimality bound. An optimality-driven perceptron obtains an $O(1/T)$ average mistake rate under separability, with a dimension-dependent finite mistake bound [2301.11260].

### FTRL, ONS, and gap-dependent methods

FTRL applied to the Fenchel–Young suboptimality loss provides a general $O(\sqrt T)$ regret guarantee. If every agent problem has a positive objective-value gap $\Delta$, a self-bounding relation between squared prediction residuals and regret yields a horizon-independent cumulative bound of order $O(\Delta^{-2})$, despite the loss being piecewise linear and lacking strong convexity [2501.13648].

Online Newton step (ONS) methods apply exp-concave surrogate losses to the inverse-learning residuals. The resulting algorithm maintains an $n\times n$ curvature matrix rather than all historical optimality constraints. It achieves

$$
O\!\left(Bn\log\left(\frac{DKT}{Bn}\right)\right)
$$

regret with per-round complexity independent of $T$. MetaGrad extends the method to suboptimal feedback and obtains a bound of order

$$
O\!\left(
Bn\log\left(\frac{DKT}{Bn}\right)
+
\sqrt{\Delta_TBn\log\left(\frac{DKT}{Bn}\right)}
\right),
$$

where $\Delta_T$ is cumulative suboptimality of the observed actions [2501.14349].

### Small-gradient skipping and finite mistakes

Small-Gradient Skipping (SGS) exploits the fact that correct recommendations yield zero residual and therefore permit a zero subgradient. Updates occur only on mistake rounds. Under a uniform separation margin $\gamma$, every mistake produces at least $\gamma$ progress relative to a separating parameter. SGS-OGD gives a finite mistake bound with quadratic dependence on $1/\gamma$, while SGS-ONS improves the dependence to roughly $\gamma^{-1}\log(1/\gamma)$.

For integer linear programs, integrality supplies positive margins under appropriate objective domains. For M-convex and $M^\natural$-convex action sets, exchange structures provide polynomial margin bounds and yield horizon-independent regret. In the simplex case, ONS and growing-grid SGS-MetaGrad attain bounds of order $O(Ld\log(2dL))$, without the center-of-gravity computation used by earlier geometric methods [2609.09809].

## 4. Regret theory and geometric structure

The main regret regimes can be distinguished by their assumptions and computational requirements.

| Setting | Principal guarantee | Main structural requirement |
|---|---:|---|
| General online inverse optimization | $O(1/\sqrt T)$ average regret | Bounded domains and feasible actions |
| Oracle-based MWU or OGD | $O(\sqrt T)$ cumulative regret | Linear optimization oracle and bounded diameter |
| Gap-dependent FTRL | Horizon-independent cumulative bound | Positive objective-value gap |
| ONS | $O(n\log T)$ cumulative regret | Exp-concave surrogate and bounded geometry |
| MOM online learning | $O(1/\sqrt T)$ average suboptimality | Nondegenerate LP bases and bounded conditioning |
| M-convex center-of-gravity method | $O(d\log d)$ regret | M-convex exchange structure |
| SGS for M-convex or ILP settings | Horizon-independent polynomial regret | Uniform separation margin |
| Variable-metric proper algorithm | $O(d)$ regret | Bounded objective domain and action sets |
| Multiscale matrix multiplicative weights | $O(\sqrt d)$ expected regret | Randomization, exact optimal feedback, oracle access |

The geometric methods maintain a region of objective vectors consistent with observed actions. For M-convex sets, an observed action implies coordinate-order constraints of the form

$$
w(i)\geq w(j)
$$

whenever the exchange $x-e_i+e_j$ is feasible. The center of gravity of the remaining normalized objective region produces a separating cut whenever the recommendation is wrong. Each mistake removes a constant fraction of the region, while the region retains an order simplex of volume $1/d!$. This yields an $O(d\log d)$ bound. Under up to $C$ corrupted observations, cycle detection and restarting give $O((C+1)d\log d)$ regret [2602.01682].

A later proper variable-metric method reduces the general horizon-independent bound to $O(d)$. It maintains a positive-definite matrix $H_t$ and query point $p_t$, with self-normalized rank-one updates

$$
s_t=\sqrt{g_t^\top H_t^{-1}g_t},
$$

$$
H_{t+1}
=
H_t+\frac{\tau}{s_t}g_tg_t^\top,
$$

and

$$
p_{t+1}
=
p_t-\alpha H_{t+1}^{-1}g_t.
$$

The trace-power potential $\operatorname{tr}(H_t^{-1/2})$ has a finite initial budget of order $d$, unlike the $\log\det H_t$ potential. This produces an $O(d)$ deterministic, proper, horizon-uniform regret bound with $O(d^2)$ matrix-update complexity and one forward optimization oracle call per round [2609.13440].

An improper randomized method achieves the dimension-optimal expected bound $O(\sqrt d)$ for arbitrary compact action sets in the Euclidean unit ball. It uses multiscale matrix multiplicative weights over polynomial feature spaces and recommends from a distribution obtained by a linear program. The guarantee is uniform in the horizon and matches an $\Omega(\sqrt d)$ lower bound up to constants. Its implementation, however, requires $(dT)^{O(d)}$ oracle calls per round and is not known to be polynomial in dimension and input length [2609.26978].

These results expose a central computational-statistical tradeoff. Proper algorithms preserve the interpretation of each recommendation as optimal for a predicted objective, whereas improper methods can obtain stronger dimension dependence. Geometric and matrix-based methods avoid storing all historical constraints, but may require difficult projections, matrix operations, center-of-gravity approximations, or large feature spaces.

## 5. Ensemble, stability, and adaptive inverse learning

Not all sequential inverse optimization uses a single observation per round. Ensemble generalized inverse linear optimization infers one cost vector from multiple decisions associated with the same feasible region. Observations may be feasible or infeasible, optimal or suboptimal, and generated by heuristics rather than exact optimization. The generalized model introduces perturbations $\epsilon_q$ and enforces dual feasibility and strong duality for perturbed observations.

Three loss families are unified:

- **Absolute duality-gap loss**, which minimizes aggregate absolute objective discrepancies.
- **Relative duality-gap loss**, which normalizes errors by the optimal objective value.
- **Decision-space loss**, which requires perturbed observations to be feasible optimal solutions and minimizes physical perturbation distances.

For feasible observations, the absolute-gap ensemble reduces analytically to a single-point model at the centroid of the observations. Exact solution methods include LP enumeration, facet-projection formulations, and convex projection problems. The framework also defines an ensemble coefficient of complementarity $\rho(\mathcal X)$, analogous to $R^2$, which compares fitted aggregate loss with facet-normal baseline loss. A larger $\rho$ indicates better scale-free fit, although it need not predict out-of-sample or clinical performance [1804.04576].

The radiation-therapy application uses eight predicted dose distributions as ensemble observations and infers one objective-weight vector for a consensus treatment plan. The relative-gap model achieved higher clinical-criteria satisfaction than the absolute-gap ensemble and several baselines in the reported experiments. The framework is batch in its formal development, but observations can be added sequentially, with centroid statistics in feasible-data cases and warm-started repeated LP solves. No online regret, stability, or consistency theorem is established.

Quantile inverse optimization addresses instability and outliers by requiring only a fraction $\theta$ of observations to be within perturbation threshold $\tau$ of optimality. The model chooses a subset $S$ with

$$
|S|\geq\theta K
$$

and seeks a cost vector for which observations in $S$ can be made optimal. Decreasing $\theta$ weakly enlarges the admissible objective set and cannot reduce inverse stability. Mixed-integer formulations encode selected observations and compatible facets, while biclique formulations identify jointly supported objective cones. An incremental online heuristic maintains a growing observation–facet incidence matrix and updates the current cone using newly arriving batches. It is a stability-oriented heuristic rather than a general regret method [1908.02376].

Data-driven inverse linear optimization provides another sequential decomposition approach when experiments have different feasible regions but share one cost vector. It separates vertex denoising from cost estimation, maintains intersections of polyhedral cones, and resolves the joint mixed-integer problem only when a new experiment makes the accumulated vertex projections incompatible. Adaptive sampling selects the next feasible-region instance to maximally reduce the current admissible cost cone. The resulting procedures reduce computation and sampling effort in customer-preference and production-planning applications, but do not establish statistical consistency or online regret guarantees [2009.07961].

Inverse Learning extends this perspective by jointly selecting a recovered objective and an optimal solution close to observed feasible and infeasible decisions. The framework distinguishes relevant, preferred, and trivial constraints and introduces Goal-integrated Inverse Learning, which controls the number and identity of constraints binding at the learned solution. Increasing the binding-constraint count generally moves the solution farther from observed behavior while satisfying more prescribed goals. Its adaptive $GIL_r$ sequence is sequential over increasingly strong goals, not temporal online learning, and the paper does not provide temporal regret or streaming guarantees [2011.03038].

## 6. Beyond objective learning: constraints and inverse optimal control

Objective learning assumes that the feasible region is known. Several approaches instead learn constraints or feasible regions.

Learning feasible regions with known linear objectives represents the learned feasible set as

$$
\mathcal X_{\theta(s)}
=
\left\{
A_{\theta(s)}z+b_{\theta(s)}:z\in\mathcal Z
\right\},
$$

where $\mathcal Z$ is a fixed primitive set and the transformation depends affinely on the observed signal. Simplex primitives represent switching and piecewise policies, Euclidean balls yield ellipsoidal regions, boxes represent bounded flows, and binary simplexes encode finite policy choices. The latent product $A_{\theta(s)}z$ makes training nonconvex. Customized block coordinate descent, adaptive smoothing, convex restrictions, and mixed-integer formulations are used to solve the resulting models. The formulation is primarily batch; stochastic updates and forgetting-factor objectives are proposed as possible online extensions, but no streaming convergence or regret theory is established [2505.15025].

Online inverse optimal control applies streaming inverse learning to discrete-time systems with active control constraints. The objective is linear in unknown basis-function coefficients:

$$
V(x_{0:K},u_{0:K},\theta)
=
\sum_{k=0}^K\theta^\prime L_k(x_k,u_k).
$$

The method uses the discrete-time minimum principle. Inactive controls supply equality stationarity conditions, while active controls contribute no equality stationarity equation. Costates are eliminated through a forward recursion, yielding a fixed-dimensional parameter vector

$$
\alpha=
\begin{bmatrix}
\theta\\
\lambda_0
\end{bmatrix}.
$$

The information matrix is updated as

$$
\mathcal Q_k
=
\mathcal Q_{k-1}
+
(F_k\mathcal G_k)^\prime(F_k\mathcal G_k)
$$

at inactive times and remains unchanged at active times, while the state–costate transition matrix is always propagated. A fixed-coordinate normalization removes positive-scale ambiguity. When the reduced information matrix has full rank and the trajectory is truly optimal for the specified model, the objective parameters are uniquely recovered under the normalization. The stored matrices have fixed dimension, independent of the trajectory horizon [2005.06153].

The control-specific method differs from generic online inverse linear optimization because its construction depends on Hamiltonians, costates, invertible state Jacobians, and minimum-principle stationarity. Its generic component is recursive least squares over a fixed-dimensional parameter vector. It illustrates how domain-specific optimality conditions can produce online inverse estimators that are substantially more memory-efficient than batch KKT formulations.

## 7. Applications, limitations, and research directions

OILO has been evaluated in dependency parsing, word alignment, shortest paths, matching, reviewer assignment, customer preference learning, knapsack, transshipment, production planning, diet recommendation, radiation therapy, power-system operation, and inverse optimal control. Across these domains, the forward optimization structure is not merely a computational subroutine: it supplies the optimality conditions, separation constraints, exchange relations, or oracle feedback used by the inverse learner.

The principal limitations are structural.

**Identifiability** is often impossible without additional assumptions. Objective scaling, normal-cone equivalence, tied optima, insufficient excitation, and redundant features can leave multiple parameters behaviorally indistinguishable. Many guarantees therefore concern decision regret or explanatory loss rather than parameter recovery.

**Regularity assumptions** vary substantially. Strong-convexity analyses do not automatically cover generic LPs. MOM requires nondegenerate bases and bounded inverse basis conditioning. Finite-mistake results require a uniform separation margin. M-convex guarantees require discrete exchange structure. Constraint-learning models require additional assumptions on feasible-region variation and latent representations.

**Computational complexity** remains a central issue. KKT reformulations may produce MISOCPs or mixed-integer quadratic programs. Quantile, vertex-selection, biclique, and feasible-region-learning models can be NP-hard or nonconvex. Exact center-of-gravity computation is $\#P$-hard. Improper multiscale methods can achieve optimal dimension dependence while requiring exponentially many oracle calls. Polynomial-time forward optimization does not imply polynomial-time inverse updating.

**Noise and suboptimality** require careful distinction. Some methods explicitly permit infeasible, noisy, heuristic, or suboptimal observations; others require exact optimal feedback. Quantile methods tolerate outliers by discarding a fraction of observations, while MetaGrad provides a suboptimality-dependent bound. Corruption-robust methods detect inconsistency through directed cycles for M-convex action sets. Robustness to arbitrary corruption in general action sets remains an active problem.

**Online adaptation** is not synonymous with repeated batch fitting. A genuine online method may be memoryless, as in implicit KKT updates; maintain sufficient statistics, as in inverse optimal control; retain a growing graph, as in quantile inverse optimization; or maintain a curvature matrix, as in ONS and variable-metric methods. Batch frameworks such as ensemble inverse optimization, Inverse Learning, and feasible-region learning provide useful components but do not automatically supply online regret or convergence guarantees.

Current research directions include closing the gap between $\Omega(d)$ and $O(d)$ or $O(d\log d)$ regret in various computational models, obtaining polynomial-time $O(\sqrt d)$ methods, improving corruption and suboptimal-feedback guarantees, extending finite-regret results beyond M-convex action sets, learning objectives and feasible regions jointly, handling partial or bandit feedback, developing statistical confidence regions, and designing scalable algorithms for nonstationary environments. The central methodological theme is the use of forward optimization structure—dual feasibility, reduced costs, exchange graphs, normal cones, Fenchel–Young losses, minimum-principle recursions, and adaptive matrix potentials—to convert sequential observations into computationally tractable inverse updates.

Source: https://www.emergentmind.com/topics/online-inverse-linear-optimization