---
title: Nested Extremum Seeking (nES)
url: https://www.emergentmind.com/topics/nested-extremum-seeking-nes
type: topic
---

# Nested Extremum Seeking (nES)

Searching arXiv for recent and foundational papers on nested extremum seeking and Lie-bracket ES.
Search query: nested extremum seeking Lie bracket averaging Stackelberg equilibrium extremum seeking nonholonomic systems
Nested extremum seeking (nES) is a multi-loop extremum-seeking architecture in which multiple ES mechanisms are stacked across separated time scales, so that a fast inner loop realizes a lower-level optimization or motion-generation task and a slower outer loop optimizes a higher-level objective using only measured performance values. In the recent game-theoretic formulation, nES is presented as multiple ES loops running simultaneously on interacting variables, with equilibrium selection determined by the hierarchy of gains and dither frequencies rather than by changing the closed-loop structure itself [2603.24756]. In the nonholonomic control setting, the same nested idea appears as an inner Lie-bracket-based oscillatory tracking layer wrapped by an outer model-free ES layer, yielding a rigorous multiple-time-scale extremum-seeking scheme for driftless control-affine systems [2005.11370].

## 1. Definition and architectural forms

The defining feature of nES is the nesting of ES loops in series. In the two-level case emphasized in current work, a leader loop and a follower loop evolve according to
$$
\begin{aligned}
\dot{x}_1 &= \sqrt{\alpha_1\omega_1}\cos\big(\omega_1 t + k_1 J_{\rm L}(x_1, x_2)\big),\\
\dot{x}_2 &= \sqrt{\alpha_2\omega_2}\cos\big(\omega_2 t + k_2 J_{\rm F}(x_1, x_2)\big),
\end{aligned}
$$
with distinct dither frequencies and positive gains [2603.24756]. This representation is model-free in the usual ES sense: the costs are measured in real time, but analytic gradients are not required.

A second, structurally different but conceptually equivalent realization appears for nonholonomic systems. There, the plant is a driftless control-affine system
$$
\dot x = \sum_{i=1}^m u_i f_i(x), \qquad y = J(x),
$$
with \(m<n\), and the nested design introduces an auxiliary reference state \(\xi\in\mathbb{R}^n\) through
$$
\dot\xi = g(y,t), \qquad u_i = \varphi_i^\varepsilon(x,\xi,t).
$$
The inner layer uses model information and Lie-bracket approximations to make \(x(t)\) track \(\xi(t)\); the outer layer is a model-free ES law driven only by the scalar output \(y=J(x)\) [2005.11370].

Two architectural motifs recur across these formulations. First, the inner loop is faster than the outer loop. Second, the fast oscillations are not merely probing signals but are designed so that their averaged effect induces a useful slow vector field. In the nonholonomic setting, that slow vector field is a stabilizing or tracking dynamics in state space; in the game-theoretic setting, it is a gradient or best-response-like dynamics. This suggests that nES is best understood not as a single algorithmic template but as a hierarchy of oscillatory feedback layers whose averaged interactions realize a prescribed optimization geometry.

## 2. Lie-bracket and averaging foundations

The principal analytical foundation of nES is the Lie-bracket interpretation of extremum seeking. For an input-affine system with fast periodic perturbations,
$$
\dot{x} = b_0(t,x) + \sum_{i=1}^{m} b_i(t,x)\sqrt{\omega}\, u_i(t,\omega t),
$$
the associated Lie-bracket system is
$$
\dot{z} = b_0(t,z) + \sum_{i<j} [b_i,b_j](t,z)\,\nu_{ji}(t),
$$
where \(\nu_{ji}(t)\) depends only on the dithers [1109.6129]. The central conclusion is that, for sufficiently large dither frequency, trajectories of the true ES system are approximated on finite horizons by trajectories of this bracket system, and uniform asymptotic stability of the bracket system implies practical uniform asymptotic stability of the original ES dynamics [1109.6129].

This viewpoint is especially important for nES because each loop can be designed so that its effective slow dynamics is itself the target of a higher-level loop. In the simplest scalar example,
$$
\dot{x} = \alpha\sqrt{\omega}\cos(\omega t)+ f(x)\sqrt{\omega}\sin(\omega t),
$$
the Lie-bracket reduction yields
$$
\dot{z} = \frac{\alpha}{2}\nabla f(z),
$$
making the optimizing behavior explicit [1109.6129]. In nested designs, the same mechanism is applied iteratively: fast dithers generate an averaged drift, and slower loops act on the state produced by that averaged drift.

Recent two-level Stackelberg analysis makes this iterative structure explicit through a three-step reduction: fast two-dimensional Lie-bracket averaging for the follower loop, singular perturbation to exploit fast follower adaptation, and slow one-dimensional Lie-bracket averaging for the leader loop [2603.24756]. In the nonholonomic setting, the interaction between layers is analyzed with averaging and Chen–Fliess expansions, and the remainder terms are controlled through constructive separation conditions on the time-scale parameters \(\varepsilon\) and \(\mu\) [2005.11370]. Taken together, these results place nES within the broader theory of oscillatory control, averaged systems, and singular perturbations.

## 3. Nonholonomic nested extremum seeking

A prototypical nES construction for control systems is developed for nonlinear driftless control-affine systems satisfying a one-step bracket generating condition,
$$
{\rm span}\big\{ f_i(x), [f_{j_1},f_{j_2}](x)\big\}=\mathbb{R}^n,
$$
with an invertible matrix \(\mathcal F(x)\) built from selected vector fields and first-order Lie brackets [2005.11370]. The cost \(J\) is assumed twice continuously differentiable and strongly convex on \(D\), with unique minimizer \(x^*\).

The nested closed-loop system contains two explicit time-scale parameters. The inner stabilizing layer uses oscillatory controls \(u_i=\varphi_i^\varepsilon(x,\xi,t)\) with fast scale \(\varepsilon>0\), while the outer ES layer evolves the virtual reference \(\xi\) on the slower scale \(\mu>0\), under the design condition
$$
\varepsilon < \mu.
$$
The inner control contains a constant component that generates motion along directly actuated vector fields and a high-frequency sinusoidal component, with amplitudes scaling like \(\sqrt{1/\varepsilon}\) and frequencies proportional to \(1/\varepsilon\), that generates averaged motion along Lie brackets \([f_{i_1},f_{i_2}]\) [2005.11370]. Its coefficients are chosen through
$$
a(x,\xi)=-\gamma_1\mathcal F^{-1}(x)(x-\xi),
$$
so that the averaged inner dynamics approximates \(\dot x \approx -\gamma_1(x-\xi)\).

The outer loop has the form
$$
\dot \xi = \sum_{j=1}^{2n} g_j(y) v_j^\mu(t)e_j,\qquad y=J(x),
$$
with generating functions satisfying
$$
[g_j(z),g_{j+n}(z)] = -\gamma_2,\qquad j=1,\dots,n,
$$
and dither amplitudes proportional to \(1/\sqrt{\mu}\) with frequencies proportional to \(1/\mu\) [2005.11370]. After averaging, the \(\xi\)-dynamics becomes
$$
\dot \xi \approx -\gamma_2\nabla J(\xi),
$$
so the outer loop behaves like gradient descent in the virtual state.

The analysis yields exponential convergence of \(x(t)\) to an arbitrarily small neighborhood of \(x^*\) under suitable choices of \(\mu\), \(\varepsilon\), and \(\gamma_1\). More precisely, for initial conditions in a sufficiently small ball around \(x^*\),
$$
\|x(t)-x^*\| \le \beta\|x^0-x^*\|e^{-\lambda t} + \rho,\qquad \forall t\ge 0,
$$
for some \(\beta,\lambda>0\), with residual \(\rho\) tunable through the time-scale parameters [2005.11370].

The Brockett integrator,
$$
\dot x_1=u_1,\qquad \dot x_2=u_2,\qquad \dot x_3=u_1x_2-u_2x_1,
$$
serves as the numerical example. Its vector fields satisfy
$$
[f_1,f_2](x)=(0,0,-2)^\top,\qquad {\rm span}\{f_1,f_2,[f_1,f_2]\}=\mathbb{R}^3,
$$
and the chosen objective is \(J(x)=\|x\|^2\) [2005.11370]. Two generating-function choices are illustrated: the classical \(g_1(z)=z,\ g_2(z)=1\), and a pair that vanishes at \(z=0\), which reduces steady-state oscillations while preserving the nested architecture.

## 4. Nash and Stackelberg regimes

In game-theoretic nES, the same two-loop dynamics can realize either Nash or Stackelberg behavior. A Nash equilibrium \((x_1^{\rm N},x_2^{\rm N})\) satisfies
$$
\partial_{x_1} J_{\rm L}(x_1^{\rm N}, x_2^{\rm N}) = 0,\qquad
\partial_{x_2} J_{\rm F}(x_1^{\rm N}, x_2^{\rm N}) = 0,
$$
whereas a Stackelberg equilibrium is defined through the follower best-response map
$$
h(x_1)=\arg\min_{x_2} J_{\rm F}(x_1,x_2)
$$
and the reduced leader cost
$$
\tilde J_{\rm L}(x_1)=J_{\rm L}(x_1,h(x_1)).
$$
The Stackelberg point \((x_1^{\rm S},x_2^{\rm S})\) then satisfies
$$
x_1^{\rm S}\in\arg\min_{x_1}\tilde J_{\rm L}(x_1),\qquad x_2^{\rm S}=h(x_1^{\rm S}) \ [2603.24756].
$$

The crucial result is that equilibrium selection depends on parameter hierarchy, not on changing the loop structure. Stackelberg behavior is obtained by enforcing the four-time-scale ordering
$$
\alpha_1 k_1 \ll \omega_1 \ll \alpha_2 k_2 \ll \omega_2,
$$
so that follower dithering is fastest, follower adaptation is next, leader dithering is slower, and leader adaptation is slowest [2603.24756]. Under assumptions \(J_{\rm L}\in C^2(\mathbb{R}^2)\), \(J_{\rm F}\in C^3(\mathbb{R}^2)\), uniqueness and \(C^1\)-regularity of \(h\), and strong convexity of both \(J_{\rm F}\) in \(x_2\) and \(\tilde J_{\rm L}\), the original nES system is \(\upsilon\)-SPUAS to the unique Stackelberg equilibrium for sufficiently large \(\omega_1\), sufficiently small \(\varepsilon=1/(\alpha_2 k_2)\), and sufficiently large \(\omega_2\) [2603.24756].

The quadratic example
$$
J_{\rm L}(x_1,x_2)=\tfrac12 x_1^2+2x_1x_2,\qquad
J_{\rm F}(x_1,x_2)=\tfrac12(x_2-2x_1+1.5)^2
$$
makes the distinction explicit. The follower best response is \(h(x_1)=2x_1-1.5\); the Nash equilibrium is
$$
(x_1^{\rm N},x_2^{\rm N})=(0.6,-0.3),
$$
whereas the Stackelberg equilibrium is
$$
(x_1^{\rm S},x_2^{\rm S})=\left(\tfrac13,-\tfrac56\right)
$$
[2603.24756]. Simulations reported there show convergence to a neighborhood of the Nash point under comparable time scales and to a neighborhood of the Stackelberg point under hierarchical scaling.

The Fish War game provides a second illustration. With
$$
\begin{aligned}
J_{\rm L}(u,v) &= -\log u - \beta_L \log\big(x-u-v^{\mu_L}\big)^\tau,\\
J_{\rm F}(u,v) &= -\log v - \beta_F \log\big(x-v-u^{\mu_F}\big)^\tau,
\end{aligned}
$$
and parameters
$$
(\tau,\mu_L,\mu_F,\beta_L,\beta_F,x)=(0.2852,1.1,1.2,0.8,0.48,1.259),
$$
the reported equilibria are \((u^{\rm N},v^{\rm N})=(0.3,0.9)\) and \((u^{\rm S},v^{\rm S})=(1.19426,0.01896)\) [2603.24756]. Although the assumptions of the theorem do not hold globally for this example, the same scaling mechanism is reported to work locally.

## 5. Learning dynamics and effective objectives

A recurrent misconception is that ES, and therefore nES, directly optimizes the original unknown objective. A more precise statement is that perturbation-based ES recovers an averaged gradient, and the induced learning dynamics can be interpreted as gradient descent on an effective objective \(L_\omega\), not necessarily on the original \(F\) [1809.04532].

For the scalar ES system
$$
\dot{x}(t)=g_1(F(x(t)))u_1(t)+g_2(F(x(t)))u_2(t),
$$
with \(T\)-periodic dithers and Lie-bracket relation
$$
[g_1,g_2](F)=-g_0(F),
$$
needle-variation analysis yields a recursion at period sampling times \(T_k=(k-1)T\),
$$
x(T_{k+1}) = x(T_k) + \nabla L_\omega(x(T_k)),
$$
where \(\nabla L_\omega\) is expressed as a double integral involving the nominal trajectory, the dither signals, the state-transition matrix of the variational equation, and the derivative of \(F\) [1809.04532]. The recovered gradient is therefore nonlocal and averaged over a period.

The multidimensional extension uses sequential dithers so that each coordinate receives a component-wise averaged partial derivative with only higher-order cross-coupling [1809.04532]. For nested architectures, the direct implication is that each loop optimizes a smoothed version of the function it sees. This suggests that, under sufficient time-scale separation, an outer loop acts on a cost landscape already modified by the inner loop’s effective objective. A plausible implication is that the nested steady state is generally the minimizer of a composed effective cost rather than of the original multilevel objective.

This effective-objective viewpoint also clarifies why ES can sometimes traverse local extrema of the original function. Because the recovered gradient is averaged over a finite time window, local features can be evened out in the learning dynamics [1809.04532]. In nES, the same smoothing may appear at multiple layers.

## 6. Assumptions, limitations, and research directions

Existing rigorous nES analyses rely on strong regularity and separation assumptions. The Lie-bracket approximation framework assumes \(C^2\) vector fields, bounded derivatives on compact sets, and periodic zero-mean perturbations [1109.6129]. The nonholonomic construction assumes a one-step bracket generating condition, strong convexity and smoothness bounds for \(J\), bounded inverse \(\mathcal F^{-1}(x)\), and a noise-free setting with access to the full state \(x\) for the inner control law [2005.11370]. The Stackelberg result is currently proved for two players with scalar decision variables under global strong convexity and smoothness assumptions on follower and reduced leader costs [2603.24756].

Several limitations are explicit. Higher-order Lie brackets are not treated in the nonholonomic scheme, although their extension is identified as future work [2005.11370]. The Lie-bracket ES theory provides only practical, not exact, asymptotic convergence for the perturbed system [1109.6129]. The learning-dynamics characterization does not furnish a closed-form formula for \(L_\omega\) in general, and it does not address measurement noise or saturation effects [1809.04532]. The Stackelberg proof does not yet cover \(n\)-nested hierarchies or higher-dimensional decision variables [2603.24756].

At the same time, the present literature points to several coherent directions. One is deeper hierarchy: multiple nested small parameters, each supporting a separate averaging or singular-perturbation reduction. Another is structural generalization: nonholonomic systems requiring higher-order brackets, broader classes of cost functions beyond the quadratic-like regime, and multi-agent settings in which distinct rationally related frequencies suppress unwanted mixed brackets [2005.11370; 1109.6129]. Application domains already identified include power grids, networked dynamical systems, and tuning of particle accelerators, all of which naturally exhibit hierarchical or multi-time-scale organization [2603.24756].

In this sense, nES occupies an intersection of model-free optimization, geometric control, and hierarchical dynamical systems. Its technical core is not merely the use of dithers, but the deliberate construction of layered averaged dynamics whose slow limit reproduces a desired optimization principle.

Source: https://www.emergentmind.com/topics/nested-extremum-seeking-nes