Papers
Topics
Authors
Recent
Search
2000 character limit reached

Pessimistic Bilevel Optimization Explained

Updated 14 July 2026
  • Pessimistic bilevel optimization (PBO) is a formulation that models decision making by selecting the worst-case follower response among multiple lower-level optima.
  • It integrates variational analysis, set-valued mappings, and KKT-based reformulations to handle the challenges of nonsmooth and nonconvex upper-level objectives.
  • PBO is applied in robust optimization, strategic classification, and hyperparameter tuning, while its computational and theoretical aspects continue to inspire new research directions.

Pessimistic bilevel optimization (PBO) is the bilevel formulation in which the leader chooses an upper-level decision xx, the follower solves a lower-level optimization problem, and—when the follower has multiple optimal responses—the leader evaluates xx against the follower-optimal response that is worst for the leader. In the standard notation S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}, the pessimistic value function is Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}, and the upper level solves minxQΨ(x)\min_{x\in Q}\Psi(x) or, equivalently, minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y) (Som et al., 23 Oct 2025, Mordukhovich, 2019). PBO is the canonical way to make a bilevel problem well defined under lower-level nonuniqueness, and it is repeatedly interpreted as a worst-case, adversarial, or robust variant of Stackelberg decision making (Som et al., 23 Oct 2025).

1. Formal structure and pessimistic semantics

In the general framework, the leader chooses xXXx\in X\subseteq \mathcal X, the follower chooses yYy\in \mathcal Y subject to yK(x)y\in K(x), and the lower-level problem is

miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.

Its solution set is

xx0

The raw upper-level problem,

xx1

is ill-defined when xx2 is not a singleton, because the leader’s objective is not a function of xx3 alone but depends on which xx4 is selected (Som et al., 23 Oct 2025).

PBO resolves this ambiguity by replacing the multivalued upper-level payoff with the pessimistic value function

xx5

and then solving the real pessimistic bilevel problem

xx6

A global real pessimistic solution is xx7 such that xx8 for all xx9; the local notion restricts the inequality to a neighborhood of S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}0 (Som et al., 23 Oct 2025). In the notation used in variational analysis, the same construction appears as

S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}1

with the pessimistic upper level S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}2 (Mordukhovich, 2019).

A central technical point is that PBO does not require value attainment. The supremum in S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}3 may fail to be attained even when S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}4 is finite and meaningful, so replacing S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}5 by S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}6 changes the model unless attainment is known a priori (Som et al., 23 Oct 2025). In algorithmic treatments, this often manifests as an effective min–max–min structure,

S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}7

whose upper-level value function is non-smooth and typically non-convex (Qichao et al., 30 Sep 2025).

2. Set-valued and two-level value-function viewpoints

When the lower-level solution map is multivalued, the upper-level objective can be treated as a set-valued map

S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}8

This yields a set-valued optimization problem S(x)=argminy{f(x,y)g(x,y)0}S(x)=\arg\min_y\{f(x,y)\mid g(x,y)\le0\}9 over Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}0. In the scalar case Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}1 with the cone Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}2, pessimistic behavior is captured by Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}3-minimality: if Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}4 and Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}5, then Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}6. The core equivalence is that every global or local Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}7-minimal solution of the set-valued problem is a global or local real pessimistic solution of the bilevel problem, and conversely, if each Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}8 is closed-valued, every global or local real pessimistic solution is Ψ(x)=sup{F(x,y)yS(x)}\Psi(x)=\sup\{F(x,y)\mid y\in S(x)\}9-minimal (Som et al., 23 Oct 2025).

The relation becomes subtler when the supremum is not attained. Let minxQΨ(x)\min_{x\in Q}\Psi(x)0 be the set of global real pessimistic solutions and let

minxQΨ(x)\min_{x\in Q}\Psi(x)1

If every global PBO solution has the supremum attained, then every PBO solution is a global minxQΨ(x)\min_{x\in Q}\Psi(x)2-minimal solution. If some global PBO solutions do not attain the supremum, then those non-attaining solutions are exactly the global minxQΨ(x)\min_{x\in Q}\Psi(x)3-minimal solutions, while the attaining ones are not minxQΨ(x)\min_{x\in Q}\Psi(x)4-minimal. The local analogue has the same attainment-sensitive structure (Som et al., 23 Oct 2025).

The two-level value-function approach formulates pessimism through

minxQΨ(x)\min_{x\in Q}\Psi(x)5

with the pessimistic bilevel program minxQΨ(x)\min_{x\in Q}\Psi(x)6. In nonsmooth analysis, a useful identity is

minxQΨ(x)\min_{x\in Q}\Psi(x)7

so the pessimistic two-level value function is the negative of an optimistic one built from minxQΨ(x)\min_{x\in Q}\Psi(x)8. This permits transferring parts of the optimistic sensitivity analysis to the pessimistic case, while the active follower set is replaced by the set of worst-case follower responses minxQΨ(x)\min_{x\in Q}\Psi(x)9 (Dempe et al., 2017).

At a conceptual level, the set-valued reformulation clarifies the geometry of pessimism, but the general-case analysis concludes that the set-valued formulation “may not hold any bigger advantage than the existing optimistic and pessimistic formulation” (Som et al., 23 Oct 2025). This suggests that its main role is explanatory and comparative rather than decisively computational.

3. Variational analysis, KKT reformulations, and relaxation methods

From the variational-analysis viewpoint, PBO is the minimization of a supremum marginal function. The lower level is still a minimization problem and can be analyzed through coderivatives and the lower-level value function, but the upper level now involves

minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)0

The key difficulty is that “subdifferentiation of the supremum marginal functions is more involved,” and the implementation of such formulas in pessimistic bilevel programs is identified as a challenging issue (Mordukhovich, 2019). The nonsmooth two-level value-function framework extends this by deriving local Lipschitzness and subdifferential estimates for minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)1 under calmness and coderivative-based constraint qualifications, together with necessary optimality conditions for local minimizers of pessimistic bilevel programs with locally Lipschitz data (Dempe et al., 2017).

Under a convex lower level and the Slater constraint qualification, the follower’s optimality system can be encoded by KKT conditions. With

minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)2

the pessimistic value function can be rewritten as

minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)3

and the outer problem becomes minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)4 (Benchouk et al., 2024). This is not a standard MPCC in the usual sense, because the complementarity conditions are not part of the outer feasible set; they sit inside the inner maximization that defines the objective.

This KKT-based formulation supports relaxation methods. Five schemes are studied in direct parallel with MPCC relaxations: Scholtes, Lin–Fukushima, Kadrani–Dussault–Benchakroun, Steffensen–Ulbrich, and Kanzow–Schwartz. For the relaxed pessimistic value functions, the theory establishes convergence of global and local minimizers of the relaxed problems to minimizers of the original KKT-based pessimistic problem, together with convergence of relaxed stationary points to suitable C- or M-stationary points under semicontinuity and qualification assumptions (Benchouk et al., 2024). A dedicated analysis of the Scholtes relaxation proves that limits of sequences of global or local optimal solutions, and of stationary points, of the Scholtes-relaxed problems are corresponding solutions or stationary points of the KKT reformulation of the pessimistic bilevel program (Benchouk et al., 2021).

4. Complexity, coupling constraints, and fixed-dimension regimes

The complexity landscape of linear bilevel programs shows a sharp distinction between optimistic and pessimistic formulations in some parameter regimes. General bilevel linear programs are strongly minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)5-hard. When the number of follower constraints minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)6 is fixed, both optimistic and pessimistic formulations are polynomially solvable. In contrast, when the number of follower variables minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)7 is fixed, optimistic bilevel linear programs remain polynomially solvable, but the pessimistic problem is strongly minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)8-hard. This is stated as the first result showing that, under comparable assumptions, the pessimistic formulation is one complexity class harder than its optimistic counterpart (Ketkov et al., 19 Nov 2025).

The same note also shows that if the number of coupling constraints minxmaxyS(x)F(x,y)\min_x \max_{y\in S(x)}F(x,y)9 is fixed and either xXXx\in X\subseteq \mathcal X0 or xXXx\in X\subseteq \mathcal X1 is fixed, then the pessimistic problem is polynomially solvable. The fixed-xXXx\in X\subseteq \mathcal X2 tractability proof exploits the fact that xXXx\in X\subseteq \mathcal X3 lives in fixed dimension xXXx\in X\subseteq \mathcal X4, so the arrangement of basis-feasibility regions is combinatorially manageable; the fixed-xXXx\in X\subseteq \mathcal X5 hardness proof uses pessimistic coupling constraints to encode Maximum Independent Set (Ketkov et al., 19 Nov 2025).

For linear pessimistic problems with coupling constraints, a separate structural result refutes the common belief that these problems are inherently harder than pessimistic problems without coupling constraints. Any pessimistic linear bilevel problem with coupling constraints can be transformed into a pessimistic problem without coupling constraints that has the same set of globally optimal solutions and the same optimal value; it can also be transformed into an optimistic problem without coupling constraints with the same set of globally optimal leader decisions and the same optimal value (Henke et al., 3 Mar 2025). The paper organizes these equivalences through the chain

xXXx\in X\subseteq \mathcal X6

These results separate two issues that are often conflated. One is worst-case selection among lower-level optima, which is the defining feature of PBO; the other is the presence of coupling constraints. In linear models, the second does not create a fundamentally new global-solution class once the equivalence machinery is available (Henke et al., 3 Mar 2025).

5. Algorithmic formulations and computational methods

Current computational work on PBO spans exact duality-based reformulations, value-function-based sequential minimization, single-loop first-order methods, KKT relaxations, and direct solvers for necessary-condition systems. The dominant distinction is between methods that exploit a special lower-level structure exactly and methods that regularize or smooth the pessimistic value function.

Approach Core construction Scope
Duality-based QCQP (Bucarey et al., 2023) Exact reformulation of PBO as a non-convex QCQP Decision-focused prediction with linear lower-level LPs
BVFSM (Liu et al., 2021) Regularized value functions plus penalties/barriers Optimistic, pessimistic, and constrained BLO without lower-level convexity assumption
SiPBA (Qichao et al., 30 Sep 2025) Smooth approximation and single-loop projected ascent–descent/gradient descent Fully first-order PBO under strong concavity of xXXx\in X\subseteq \mathcal X7 and convexity of xXXx\in X\subseteq \mathcal X8
KKT relaxation methods (Benchouk et al., 2024) Scholtes, LF, KDB, SU, KS relaxations of the KKT-based pessimistic problem Smooth PBO with convex lower level and Slater condition
Scholtes-specific analysis (Benchouk et al., 2021) Scholtes relaxation for the pessimistic KKT reformulation Global/local solution and stationary-point convergence
LM-based system solver (Benfield et al., 2024) Solve a necessary optimality system xXXx\in X\subseteq \mathcal X9 Strategic classification with nonconvex, nonunique lower level

In decision-focused prediction, expected regret minimization is formulated exactly as a pessimistic bilevel optimization model in which the upper level chooses prediction parameters yYy\in \mathcal Y0 and the lower level chooses an adversarially worst optimal decision yYy\in \mathcal Y1 for each sample. When the lower-level nominal decision problem is a bounded linear program, the model can be dualized twice to yield an exact non-convex QCQP, and tractability is pursued through local search on the value function, a restricted penalized QCQP with yYy\in \mathcal Y2, valid inequalities, and nonconvex QCQP solvers (Bucarey et al., 2023).

Bilevel Value-Function-based Sequential Minimization (BVFSM) constructs a sequence of approximated single-level problems. For pessimistic BLO it uses the surrogate

yYy\in \mathcal Y3

and then minimizes yYy\in \mathcal Y4 over yYy\in \mathcal Y5. The method avoids repeated Hessian inverses and unrolled recurrent differentiation, and the asymptotic convergence analysis is established without the restrictive lower-level convexity assumption (Liu et al., 2021).

SiPBA takes a different route: it introduces a smooth approximation

yYy\in \mathcal Y6

where yYy\in \mathcal Y7 is strongly convex–concave in yYy\in \mathcal Y8. For each yYy\in \mathcal Y9, this yields a unique saddle point yK(x)y\in K(x)0, a continuously differentiable surrogate yK(x)y\in K(x)1, and an explicit gradient formula that uses only first-order derivatives. SiPBA then performs one projected ascent–descent step in yK(x)y\in K(x)2 and one projected gradient step in yK(x)y\in K(x)3 per iteration, with convergence to approximate stationarity for the original PBO under parameter schedules yK(x)y\in K(x)4, yK(x)y\in K(x)5, and decaying step sizes (Qichao et al., 30 Sep 2025).

In strategic classification under nonunique follower behavior, a different computational philosophy is used. The pessimistic two-level value function yK(x)y\in K(x)6 is combined with necessary optimality conditions

yK(x)y\in K(x)7

encoded as a system yK(x)y\in K(x)8 after setting yK(x)y\in K(x)9. A Levenberg–Marquardt method is then applied to the nonlinear least-squares objective miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.0 (Benfield et al., 2024).

6. Modeling roles and application domains

PBO is repeatedly used when the follower is adversarial or when the leader wants a worst-case guarantee across all lower-level optima. A basic modeling identity is that the robust counterpart

miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.1

is exactly a pessimistic bilevel problem with lower level miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.2 and upper-level value miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.3 (Som et al., 23 Oct 2025). This places robust optimization squarely inside the pessimistic bilevel template.

In decision-focused learning, exact empirical expected regret minimization is formulated as a pessimistic bilevel problem: miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.4 The pessimistic interpretation is essential because multiple optimal decisions under the predicted cost can lead to substantially different regret under the true cost (Bucarey et al., 2023).

In adversarial or strategic classification, implementation-time attacks are modeled as a leader–follower game in which the learner chooses classifier weights miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.5 and the adversary chooses generator parameters miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.6. The pessimistic upper level

miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.7

is used precisely because the lower level is nonconvex and admits multiple optimal solutions, so the learner must hedge against the most damaging follower-optimal response (Benfield et al., 2024).

In hyperparameter tuning, pessimistic bilevel optimization replaces the conventional optimistic assumption by

miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.8

or, in the approximate version,

miny{f(x,y)yK(x)}.\min_y\{f(x,y)\mid y\in K(x)\}.9

The computational studies on binary linear classifiers report better prediction performances than optimistic counterparts when training data are limited or testing data are perturbed (Ustun et al., 2024).

PBO also appears in stochastic and distributionally robust settings. Under moment ambiguity, a pessimistic stochastic bilevel program for sequential games is written as

xx00

and is shown to be equivalent to a generic two-stage distributionally robust stochastic program. For continuous ambiguity sets, linear decision rules lead to 0–1 SDP approximations and exact 0–1 copositive reformulations; for discrete ambiguity sets, an exact 0–1 SDP reformulation and an explicit worst-case distribution are derived (Goyal et al., 2022).

In elastic shape optimization, the leader chooses a material distribution xx01, the follower chooses loads xx02 from a compact convex set to maximize compliance, and the leader evaluates a tracking-type objective through the pessimistic mapping

xx03

The stochastic version evaluates xx04 through a convex risk measure xx05, yielding a pessimistic bilevel stochastic problem for robust shape design under manufacturing perturbations (Burtscheidt et al., 2021).

7. Limitations, controversies, and open directions

Several limitations recur throughout the literature. The first is non-attainment: without compactness or continuity, xx06 may not be attained, so a real pessimistic solution may have no corresponding xx07 with xx08. This complicates both interpretation and the relation to set-valued xx09-minimality, especially at the local level (Som et al., 23 Oct 2025). The second is analytic rather than modeling: even with Lipschitzian data, supremum marginal functions are harder to differentiate than infimum marginal functions, and this makes necessary optimality conditions for PBO more delicate than for optimistic bilevel programs (Mordukhovich, 2019).

A further limitation is that the set-valued reformulation does not automatically simplify computation. The general-case analysis explicitly states that solving for xx10-minimal solutions is essentially the same as solving scalar PBO for real pessimistic solutions, and does not yield “a genuinely new class of solutions or a significantly easier formulation” (Som et al., 23 Oct 2025). This is mirrored algorithmically: exact reformulations often produce nonconvex QCQPs or min–max problems with complementarity structure, while first-order methods such as SiPBA require strong concavity of xx11, convexity of xx12, and deterministic schedules that may not hold in broader machine-learning models (Qichao et al., 30 Sep 2025).

The relaxation literature imposes its own restrictions. The KKT-based methods assume a convex lower level, the Slater constraint qualification, and smoothness; when those hypotheses fail, the KKT reformulation is no longer equivalent to lower-level optimality, so the relaxation schemes become heuristic outside the stated regime (Benchouk et al., 2024). This explains why direct value-function approaches, nonsmooth variational analysis, and set-valued formulations remain active alongside MPCC-type methods rather than being displaced by them.

Open directions are stated explicitly. They include existence and regularity of PBO solutions under weaker assumptions, refined solution concepts that require worst-case payoff to be attained or approximately attained, algorithmic development that leverages set-valued structure, and multi-objective bilevel extensions requiring vector-space notions of worst-case behavior (Som et al., 23 Oct 2025). The variational literature adds efficient evaluation of symmetric subdifferentials for infimum and supremum marginal functions, relaxation of partial calmness requirements, and stronger algorithmic exploitation of necessary optimality conditions (Mordukhovich, 2019). On the algorithmic side, the current first-order analysis is deterministic, and extensions to stochastic PBO, such as minibatch-gradient variants, are identified as future work (Qichao et al., 30 Sep 2025).

In aggregate, the literature presents PBO as a mature modeling principle with a still-fragmented computational theory. The central object xx13 is stable as a definition, but its tractability, regularity, and algorithmic treatment remain sharply dependent on lower-level geometry, value attainment, and the choice between scalar, set-valued, KKT-based, or smoothed representations (Som et al., 23 Oct 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Pessimistic Bilevel Optimization (PBO).