---
title: Pessimistic Bilevel Optimization Explained
url: https://www.emergentmind.com/topics/pessimistic-bilevel-optimization-pbo
type: topic
---

# Pessimistic Bilevel Optimization Explained

Pessimistic bilevel optimization (PBO) is the bilevel formulation in which the leader chooses an upper-level decision \(x\), the follower solves a lower-level optimization problem, and—when the follower has multiple optimal responses—the leader evaluates \(x\) against the follower-optimal response that is worst for the leader. In the standard notation \(S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}\), the pessimistic value function is \(\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}\), and the upper level solves \(\min_{x\in Q}\Psi(x)\) or, equivalently, \(\min_x \max_{y\in S(x)}F(x,y)\) [2510.20631][1907.06140]. PBO is the canonical way to make a bilevel problem well defined under lower-level nonuniqueness, and it is repeatedly interpreted as a worst-case, adversarial, or robust variant of Stackelberg decision making [2510.20631].

## 1. Formal structure and pessimistic semantics

In the general framework, the leader chooses \(x\in X\subseteq \mathcal X\), the follower chooses \(y\in \mathcal Y\) subject to \(y\in K(x)\), and the lower-level problem is
\[
\min_y\{f(x,y)\mid y\in K(x)\}.
\]
Its solution set is
\[
\psi(x):=\arg\min\{f(x,y)\mid y\in K(x)\}.
\]
The raw upper-level problem,
\[
\underset{x}{\text{`minimize'}}\ F(x,y)\quad \text{subject to }x\in X,\ y\in \psi(x),
\]
is ill-defined when \(\psi(x)\) is not a singleton, because the leader’s objective is not a function of \(x\) alone but depends on which \(y\in\psi(x)\) is selected [2510.20631].

PBO resolves this ambiguity by replacing the multivalued upper-level payoff with the pessimistic value function
\[
F_p(x):=\sup_{y\in \psi(x)}F(x,y),
\]
and then solving the real pessimistic bilevel problem
\[
\underset{x\in X}{\text{minimize }}F_p(x).
\]
A global real pessimistic solution is \(x^*\in X\) such that \(F_p(x)\ge F_p(x^*)\) for all \(x\in X\); the local notion restricts the inequality to a neighborhood of \(x^*\) [2510.20631]. In the notation used in variational analysis, the same construction appears as
\[
\Psi(x):=\sup\{v(x,y)\mid y\in S(x)\},
\]
with the pessimistic upper level \(\min_{x\in Q}\Psi(x)\) [1907.06140].

A central technical point is that PBO does not require value attainment. The supremum in \(F_p(x)=\sup_{y\in\psi(x)}F(x,y)\) may fail to be attained even when \(F_p(x)\) is finite and meaningful, so replacing \(\sup\) by \(\max\) changes the model unless attainment is known a priori [2510.20631]. In algorithmic treatments, this often manifests as an effective min–max–min structure,
\[
\min_x\max_y\{F(x,y)\;\text{s.t.}\; y\in \arg\min_z f(x,z)\},
\]
whose upper-level value function is non-smooth and typically non-convex [2509.26240].

## 2. Set-valued and two-level value-function viewpoints

When the lower-level solution map is multivalued, the upper-level objective can be treated as a set-valued map
\[
G(x):=F(x,\psi(x))=\{F(x,y)\mid y\in\psi(x)\}\subset \mathbb R.
\]
This yields a set-valued optimization problem \(\min G(x)\) over \(x\in X\). In the scalar case \(Z=\mathbb R\) with the cone \(C=\mathbb R_+\), pessimistic behavior is captured by \(u\)-minimality: if \(A,B\subset \mathbb R\) and \(A\le_{\mathbb R_+}^u B\), then \(\sup A\le \sup B\). The core equivalence is that every global or local \(u\)-minimal solution of the set-valued problem is a global or local real pessimistic solution of the bilevel problem, and conversely, if each \(G(x)\) is closed-valued, every global or local real pessimistic solution is \(u\)-minimal [2510.20631].

The relation becomes subtler when the supremum is not attained. Let \(\hat{\mathcal Q}\) be the set of global real pessimistic solutions and let
\[
\hat{\mathcal T}:=\{x\in \hat{\mathcal Q}\mid \exists y\in \psi(x):F_p(x)=F(x,y)\}.
\]
If every global PBO solution has the supremum attained, then every PBO solution is a global \(u\)-minimal solution. If some global PBO solutions do not attain the supremum, then those non-attaining solutions are exactly the global \(u\)-minimal solutions, while the attaining ones are not \(u\)-minimal. The local analogue has the same attainment-sensitive structure [2510.20631].

The two-level value-function approach formulates pessimism through
\[
P_p(x):=\max\{F(x,y)\mid y\in S(x)\},
\]
with the pessimistic bilevel program \(\min_{x\in X}P_p(x)\). In nonsmooth analysis, a useful identity is
\[
P_p(x)=-P_o^{(-F)}(x),
\]
so the pessimistic two-level value function is the negative of an optimistic one built from \(-F\). This permits transferring parts of the optimistic sensitivity analysis to the pessimistic case, while the active follower set is replaced by the set of worst-case follower responses \(S_p(x)\) [1711.11127].

At a conceptual level, the set-valued reformulation clarifies the geometry of pessimism, but the general-case analysis concludes that the set-valued formulation “may not hold any bigger advantage than the existing optimistic and pessimistic formulation” [2510.20631]. This suggests that its main role is explanatory and comparative rather than decisively computational.

## 3. Variational analysis, KKT reformulations, and relaxation methods

From the variational-analysis viewpoint, PBO is the minimization of a supremum marginal function. The lower level is still a minimization problem and can be analyzed through coderivatives and the lower-level value function, but the upper level now involves
\[
\Psi(x)=\sup\{v(x,y)\mid y\in S(x)\}.
\]
The key difficulty is that “subdifferentiation of the supremum marginal functions is more involved,” and the implementation of such formulas in pessimistic bilevel programs is identified as a challenging issue [1907.06140]. The nonsmooth two-level value-function framework extends this by deriving local Lipschitzness and subdifferential estimates for \(P_p\) under calmness and coderivative-based constraint qualifications, together with necessary optimality conditions for local minimizers of pessimistic bilevel programs with locally Lipschitz data [1711.11127].

Under a convex lower level and the Slater constraint qualification, the follower’s optimality system can be encoded by KKT conditions. With
\[
\mathcal D(x):=\{(y,u)\mid \nabla_yL(x,y,u)=0,\ u\ge0,\ g(x,y)\le0,\ u^\top g(x,y)=0\},
\]
the pessimistic value function can be rewritten as
\[
\psi_p(x):=\max_{(y,u)\in \mathcal D(x)}F(x,y),
\]
and the outer problem becomes \(\min_{x\in X}\psi_p(x)\) [2412.11416]. This is not a standard MPCC in the usual sense, because the complementarity conditions are not part of the outer feasible set; they sit inside the inner maximization that defines the objective.

This KKT-based formulation supports relaxation methods. Five schemes are studied in direct parallel with MPCC relaxations: Scholtes, Lin–Fukushima, Kadrani–Dussault–Benchakroun, Steffensen–Ulbrich, and Kanzow–Schwartz. For the relaxed pessimistic value functions, the theory establishes convergence of global and local minimizers of the relaxed problems to minimizers of the original KKT-based pessimistic problem, together with convergence of relaxed stationary points to suitable C- or M-stationary points under semicontinuity and qualification assumptions [2412.11416]. A dedicated analysis of the Scholtes relaxation proves that limits of sequences of global or local optimal solutions, and of stationary points, of the Scholtes-relaxed problems are corresponding solutions or stationary points of the KKT reformulation of the pessimistic bilevel program [2110.13755].

## 4. Complexity, coupling constraints, and fixed-dimension regimes

The complexity landscape of linear bilevel programs shows a sharp distinction between optimistic and pessimistic formulations in some parameter regimes. General bilevel linear programs are strongly \(NP\)-hard. When the number of follower constraints \(m_f\) is fixed, both optimistic and pessimistic formulations are polynomially solvable. In contrast, when the number of follower variables \(n_f\) is fixed, optimistic bilevel linear programs remain polynomially solvable, but the pessimistic problem is strongly \(NP\)-hard. This is stated as the first result showing that, under comparable assumptions, the pessimistic formulation is one complexity class harder than its optimistic counterpart [2511.15592].

The same note also shows that if the number of coupling constraints \(|J|\) is fixed and either \(m_f\) or \(n_f\) is fixed, then the pessimistic problem is polynomially solvable. The fixed-\(m_f\) tractability proof exploits the fact that \((\mathbf w,t)\) lives in fixed dimension \(m_f+1\), so the arrangement of basis-feasibility regions is combinatorially manageable; the fixed-\(n_f\) hardness proof uses pessimistic coupling constraints to encode Maximum Independent Set [2511.15592].

For linear pessimistic problems with coupling constraints, a separate structural result refutes the common belief that these problems are inherently harder than pessimistic problems without coupling constraints. Any pessimistic linear bilevel problem with coupling constraints can be transformed into a pessimistic problem without coupling constraints that has the same set of globally optimal solutions and the same optimal value; it can also be transformed into an optimistic problem without coupling constraints with the same set of globally optimal leader decisions and the same optimal value [2503.01563]. The paper organizes these equivalences through the chain
\[
\text{Pessimistic w/ coupling}\rightarrow \text{Optimistic w/ coupling}\rightarrow \text{Optimistic w/o coupling}\rightarrow \text{Pessimistic w/o coupling}.
\]

These results separate two issues that are often conflated. One is worst-case selection among lower-level optima, which is the defining feature of PBO; the other is the presence of coupling constraints. In linear models, the second does not create a fundamentally new global-solution class once the equivalence machinery is available [2503.01563].

## 5. Algorithmic formulations and computational methods

Current computational work on PBO spans exact duality-based reformulations, value-function-based sequential minimization, single-loop first-order methods, KKT relaxations, and direct solvers for necessary-condition systems. The dominant distinction is between methods that exploit a special lower-level structure exactly and methods that regularize or smooth the pessimistic value function.

| Approach | Core construction | Scope |
|---|---|---|
| Duality-based QCQP [2312.17640] | Exact reformulation of PBO as a non-convex QCQP | Decision-focused prediction with linear lower-level LPs |
| BVFSM [2110.04974] | Regularized value functions plus penalties/barriers | Optimistic, pessimistic, and constrained BLO without lower-level convexity assumption |
| SiPBA [2509.26240] | Smooth approximation and single-loop projected ascent–descent/gradient descent | Fully first-order PBO under strong concavity of \(F(x,\cdot)\) and convexity of \(f(x,\cdot)\) |
| KKT relaxation methods [2412.11416] | Scholtes, LF, KDB, SU, KS relaxations of the KKT-based pessimistic problem | Smooth PBO with convex lower level and Slater condition |
| Scholtes-specific analysis [2110.13755] | Scholtes relaxation for the pessimistic KKT reformulation | Global/local solution and stationary-point convergence |
| LM-based system solver [2410.20284] | Solve a necessary optimality system \(\Phi(w,\theta,\zeta)=0\) | Strategic classification with nonconvex, nonunique lower level |

In decision-focused prediction, expected regret minimization is formulated exactly as a pessimistic bilevel optimization model in which the upper level chooses prediction parameters \(\omega\) and the lower level chooses an adversarially worst optimal decision \(v^i\) for each sample. When the lower-level nominal decision problem is a bounded linear program, the model can be dualized twice to yield an exact non-convex QCQP, and tractability is pursued through local search on the value function, a restricted penalized QCQP with \(\gamma^i=\kappa\), valid inequalities, and nonconvex QCQP solvers [2312.17640].

Bilevel Value-Function-based Sequential Minimization (BVFSM) constructs a sequence of approximated single-level problems. For pessimistic BLO it uses the surrogate
\[
\varphi^p_{\mu,\theta,\sigma}(x)=\max_y\left\{F(x,y)-\PP_\sigma(f(x,y)-f^*_\mu(x))-\frac{\theta}{2}\|y\|^2\right\},
\]
and then minimizes \(\varphi^p_{\mu,\theta,\sigma}(x)\) over \(x\in X\). The method avoids repeated Hessian inverses and unrolled recurrent differentiation, and the asymptotic convergence analysis is established without the restrictive lower-level convexity assumption [2110.04974].

SiPBA takes a different route: it introduces a smooth approximation
\[
\phi_{\rho,\sigma}(x)=\min_{z\in Y}\max_{y\in Y}\psi_{\rho,\sigma}(x,y,z),
\]
where \(\psi_{\rho,\sigma}\) is strongly convex–concave in \((z,y)\). For each \(x\), this yields a unique saddle point \((y^*_{\rho,\sigma}(x),z^*_{\rho,\sigma}(x))\), a continuously differentiable surrogate \(\phi_{\rho,\sigma}\), and an explicit gradient formula that uses only first-order derivatives. SiPBA then performs one projected ascent–descent step in \((y,z)\) and one projected gradient step in \(x\) per iteration, with convergence to approximate stationarity for the original PBO under parameter schedules \(\rho_k\to\infty\), \(\sigma_k\to0\), and decaying step sizes [2509.26240].

In strategic classification under nonunique follower behavior, a different computational philosophy is used. The pessimistic two-level value function \(\varphi_p(w)=\max_{\theta\in S(w)}F(w,\theta)\) is combined with necessary optimality conditions
\[
\nabla_w F(w,\theta)=0,\qquad \nabla_\theta F(w,\theta)-\lambda\nabla_\theta f(w,\theta)=0,\qquad \nabla_\theta f(w,\theta)=0,
\]
encoded as a system \(\Phi(w,\theta,\zeta)=0\) after setting \(\lambda=\zeta^2\). A Levenberg–Marquardt method is then applied to the nonlinear least-squares objective \(\|\Phi(w,\theta,\zeta)\|^2\) [2410.20284].

## 6. Modeling roles and application domains

PBO is repeatedly used when the follower is adversarial or when the leader wants a worst-case guarantee across all lower-level optima. A basic modeling identity is that the robust counterpart
\[
\min_{x\in S}\sup_{\xi\in U}\phi(x,\xi)
\]
is exactly a pessimistic bilevel problem with lower level \(\arg\min_\xi -\phi(x,\xi)\) and upper-level value \(F_p(x)=\sup_{\xi\in U}\phi(x,\xi)\) [2510.20631]. This places robust optimization squarely inside the pessimistic bilevel template.

In decision-focused learning, exact empirical expected regret minimization is formulated as a pessimistic bilevel problem:
\[
\min_{\omega}\max_{v^i\in V^*(\hat c^i(\omega))}\frac1N\sum_{i=1}^N\big[(c^i)^\top v^i-z^*(c^i)\big].
\]
The pessimistic interpretation is essential because multiple optimal decisions under the predicted cost can lead to substantially different regret under the true cost [2312.17640].

In adversarial or strategic classification, implementation-time attacks are modeled as a leader–follower game in which the learner chooses classifier weights \(w\) and the adversary chooses generator parameters \(\theta\). The pessimistic upper level
\[
\min_w\max_{\theta\in S(w)}F(w,\theta)
\]
is used precisely because the lower level is nonconvex and admits multiple optimal solutions, so the learner must hedge against the most damaging follower-optimal response [2410.20284].

In hyperparameter tuning, pessimistic bilevel optimization replaces the conventional optimistic assumption by
\[
P_p^*=\min_{\lambda\in\Lambda}\max_{\theta\in\Psi(\lambda)}L_{S_{\mathrm{val}}}(f(\theta)),
\]
or, in the approximate version,
\[
P_p^\varepsilon=\min_\lambda \max_{\theta:\,L_{S_{\mathrm{train}}}(f(\theta))\le (1+\varepsilon)L_{S_{\mathrm{train}}}(f(\bar\theta))}L_{S_{\mathrm{val}}}(f(\theta)).
\]
The computational studies on binary linear classifiers report better prediction performances than optimistic counterparts when training data are limited or testing data are perturbed [2412.03666].

PBO also appears in stochastic and distributionally robust settings. Under moment ambiguity, a pessimistic stochastic bilevel program for sequential games is written as
\[
\min_{x\in\mathcal X}\Big\{w^\top x+\sup_{F\in\mathcal D}\mathbb E_F\big[\max_{y\in \Omega(x,\xi)}v(\xi)^\top y\big]\Big\},
\]
and is shown to be equivalent to a generic two-stage distributionally robust stochastic program. For continuous ambiguity sets, linear decision rules lead to 0–1 SDP approximations and exact 0–1 copositive reformulations; for discrete ambiguity sets, an exact 0–1 SDP reformulation and an explicit worst-case distribution are derived [2206.03531].

In elastic shape optimization, the leader chooses a material distribution \(u\), the follower chooses loads \(f\) from a compact convex set to maximize compliance, and the leader evaluates a tracking-type objective through the pessimistic mapping
\[
\Phi[u]=\max_{f\in\Psi[u]}J[u,f].
\]
The stochastic version evaluates \(\Phi[u\odot \Upsilon]\) through a convex risk measure \(\mathcal R\), yielding a pessimistic bilevel stochastic problem for robust shape design under manufacturing perturbations [2103.02281].

## 7. Limitations, controversies, and open directions

Several limitations recur throughout the literature. The first is non-attainment: without compactness or continuity, \(\sup_{y\in\psi(x)}F(x,y)\) may not be attained, so a real pessimistic solution may have no corresponding \((x,y)\) with \(F(x,y)=F_p(x)\). This complicates both interpretation and the relation to set-valued \(u\)-minimality, especially at the local level [2510.20631]. The second is analytic rather than modeling: even with Lipschitzian data, supremum marginal functions are harder to differentiate than infimum marginal functions, and this makes necessary optimality conditions for PBO more delicate than for optimistic bilevel programs [1907.06140].

A further limitation is that the set-valued reformulation does not automatically simplify computation. The general-case analysis explicitly states that solving for \(u\)-minimal solutions is essentially the same as solving scalar PBO for real pessimistic solutions, and does not yield “a genuinely new class of solutions or a significantly easier formulation” [2510.20631]. This is mirrored algorithmically: exact reformulations often produce nonconvex QCQPs or min–max problems with complementarity structure, while first-order methods such as SiPBA require strong concavity of \(F(x,\cdot)\), convexity of \(f(x,\cdot)\), and deterministic schedules that may not hold in broader machine-learning models [2509.26240].

The relaxation literature imposes its own restrictions. The KKT-based methods assume a convex lower level, the Slater constraint qualification, and smoothness; when those hypotheses fail, the KKT reformulation is no longer equivalent to lower-level optimality, so the relaxation schemes become heuristic outside the stated regime [2412.11416]. This explains why direct value-function approaches, nonsmooth variational analysis, and set-valued formulations remain active alongside MPCC-type methods rather than being displaced by them.

Open directions are stated explicitly. They include existence and regularity of PBO solutions under weaker assumptions, refined solution concepts that require worst-case payoff to be attained or approximately attained, algorithmic development that leverages set-valued structure, and multi-objective bilevel extensions requiring vector-space notions of worst-case behavior [2510.20631]. The variational literature adds efficient evaluation of symmetric subdifferentials for infimum and supremum marginal functions, relaxation of partial calmness requirements, and stronger algorithmic exploitation of necessary optimality conditions [1907.06140]. On the algorithmic side, the current first-order analysis is deterministic, and extensions to stochastic PBO, such as minibatch-gradient variants, are identified as future work [2509.26240].

In aggregate, the literature presents PBO as a mature modeling principle with a still-fragmented computational theory. The central object \(\min_x \sup_{y\in S(x)}F(x,y)\) is stable as a definition, but its tractability, regularity, and algorithmic treatment remain sharply dependent on lower-level geometry, value attainment, and the choice between scalar, set-valued, KKT-based, or smoothed representations [2510.20631].

Source: https://www.emergentmind.com/topics/pessimistic-bilevel-optimization-pbo