Papers
Topics
Authors
Recent
Search
2000 character limit reached

Unbiased Extremum Seeking

Updated 12 July 2026
  • Unbiased extremum seeking is defined as designs that achieve exact convergence to the true optimizer by eliminating steady-state dither bias through time-varying perturbation amplitudes and gains.
  • It employs coordinated mechanisms such as chirp signals, exponential decay of dithers, and Lie bracket averaging to overcome local-extrema bias and ensure uniform convergence.
  • Advanced variants extend these methods to complex settings, including systems with delays, PDE dynamics, and distributed optimization for tracking time-varying optima.

Unbiased extremum seeking is a class of extremum-seeking (ES) designs in which the closed-loop state converges to the true extremum of an unknown objective, rather than to a dither-induced neighborhood or periodic orbit around it. In recent work, the term is used most explicitly for schemes that remove steady-state oscillation bias by combining time-varying perturbation amplitudes, time-varying demodulation or adaptation gains, and averaging arguments that preserve the optimizer as the actual limiting equilibrium (Yilmaz et al., 2024, Yilmaz et al., 2024). Closely related literature treats “unbiasedness” in broader senses as well: uniformity with respect to dither amplitude or cost magnitude, practical convergence to the global minimizer despite local minima, and reduction of bias toward undesired local extrema through non-local averaging (Mimmo et al., 2022, Suttner et al., 1 Mar 2026).

1. Classical bias and the emergence of the unbiased formulation

Classical ES perturbs the decision variable with a persistent periodic dither, demodulates the measured objective, and updates the parameter estimate using the resulting gradient surrogate. In maximum power point tracking, for example, the plant input is typically written as

d(t)=d^(t)+asin⁡(ωt),d(t)=\hat d(t)+a\sin(\omega t),

with a high-pass filter and a low-pass filtered demodulation stage. Because the dither amplitude is constant, the true optimizer is not an equilibrium of the exact closed loop, and the asymptotic behavior is a periodic orbit around the extremum rather than exact convergence (Yilmaz et al., 7 Oct 2025). The same basic phenomenon appears in PDE-compensating ES: persistent perturbations and actuator dynamics imply local convergence only to a neighborhood of the optimizer, which the unbiased PDE designs were introduced to remove (Yilmaz et al., 2024).

Recent uES papers therefore define unbiasedness as exact convergence of the controlled variable to the true optimizer, with the perturbation amplitude vanishing asymptotically or by a prescribed time. In the time-varying setting, the same idea is extended from static optima to moving optima, so that the tracking error tends to zero rather than remaining of order O(a)O(a), O(ε)O(\varepsilon), or O(ω−1)O(\omega^{-1}) (Yilmaz et al., 2024). A distinct but related line of work uses the same word less literally: in non-convex scalar ES, “unbiasedness” is associated with controlled dependence on dither amplitude, attraction to the global minimizer despite local minima, and gain bounds that are independent of the absolute magnitude of the cost function (Mimmo et al., 2022).

A second source of classical bias is topological rather than asymptotic. Standard perturbation-based ES is fundamentally local: it follows local gradient information and therefore converges to nearby local extrema when the landscape is multimodal. This “local-extrema bias” motivates non-local ES constructions that optimize a smoothed objective Jˉa\bar J_a instead of the original JJ, thereby eliminating undesired critical points of the averaged objective under suitable assumptions (Suttner et al., 1 Mar 2026).

2. Core algorithmic mechanisms

The characteristic mechanism of unbiased ES is coordinated time variation in excitation and demodulation. In the exponential uES scheme for MPPT, the perturbation amplitude is scaled by

α˙(t)=−λα(t),α(t)=α0e−λ(t−t0),\dot{\alpha}(t)=-\lambda \alpha(t), \qquad \alpha(t)=\alpha_0 e^{-\lambda (t-t_0)},

and the plant input is

d(t)=d^(t)+α(t) asin⁡(ωt).d(t)=\hat d(t)+\alpha(t)\,a\sin(\omega t).

The measured power is high-pass filtered through

η˙=−ωhη+ωhP(d(t)),\dot{\eta}=-\omega_h\eta+\omega_h P(d(t)),

while demodulation uses the inverse scaling 1/α(t)1/\alpha(t), so that gradient information is preserved even as the dither vanishes (Yilmaz et al., 7 Oct 2025). The prescribed-time version replaces the exponential decay with a blow-up clock,

O(a)O(a)0

and chirp signals whose instantaneous frequency increases as O(a)O(a)1 approaches the prescribed horizon O(a)O(a)2, allowing convergence by a user-defined time (Yilmaz et al., 7 Oct 2025).

The same design logic appears in the general time-varying uES framework. There the central transformation is a scaled tracking error such as

O(a)O(a)3

or O(a)O(a)4, where O(a)O(a)5 or O(a)O(a)6 is chosen to grow polynomially, exponentially, or with prescribed-time blow-up. The perturbation amplitude decays relative to the original state coordinates, while the transformed error dynamics become asymptotically stable at the origin. This produces exact tracking of time-varying optima and allows asymptotic, exponential, or prescribed-time rates, depending on the choice of scaling and probing frequency law (Yilmaz et al., 2024).

In PDE-compensated uES, unbiasedness is achieved with the same asymmetry between additive and multiplicative channels. The additive dither decays exponentially, but the demodulation signals grow exponentially: O(a)O(a)7 Combined with a high-pass filter and PDE backstepping, this yields exponential convergence of the actual actuator state to the true optimizer, rather than convergence to a PDE-distorted neighborhood (Yilmaz et al., 2024).

A discrete-time variant for static quadratic maps with unknown large time-varying delays uses the same principle in sampled form. The injected perturbation has exponentially decaying envelope

O(a)O(a)8

and the estimator update is

O(a)O(a)9

A high-pass filter

O(ε)O(\varepsilon)0

removes the constant offset, while dither frequencies are chosen of order O(ε)O(\varepsilon)1 to handle arbitrarily large unknown bounded delays (Jbara et al., 6 Apr 2026).

3. Averaging, Lie brackets, and uniformity

The analytical foundation of unbiased ES is averaging. In the classical Lie-bracket interpretation of ES, a sinusoidally perturbed control-affine system is approximated by a Lie-bracket system whose vector field directly reveals the optimizing behavior. For the scalar prototype

O(ε)O(\varepsilon)2

the associated Lie-bracket system is

O(ε)O(\varepsilon)3

so the reduced dynamics ascend the gradient of the unknown map itself (Dürr et al., 2011). This viewpoint makes clear that any residual bias in the original ES loop is a finite-frequency effect rather than a structural displacement of the optimizer in the ideal averaged dynamics.

Recent work extends this logic globally. “Initialization-Free Lie-Bracket Extremum Seeking in O(ε)O(\varepsilon)4” develops a second-order averaging result with global practical stability implications, allowing uniform global practical asymptotic stability under conditions that include globally Lipschitz gradient O(ε)O(\varepsilon)5, rather than globally Lipschitz O(ε)O(\varepsilon)6 itself. This covers quadratic costs and certain non-smooth Lie-bracket ES dynamics, and removes the semi-global initialization dependence that dominates earlier results (Abdelgalil et al., 2024).

A complementary scalar line of work replaces Taylor-expansion analysis by Fourier analysis of the exact averaged system. For the dithered measurement O(ε)O(\varepsilon)7, the average dynamics reduce exactly to

O(ε)O(\varepsilon)8

where O(ε)O(\varepsilon)9 is the first sine Fourier coefficient of the dithered cost. Under a unique global minimizer and a quasi-convex or non-convex envelope assumption, the averaged attractor lies in O(ω−1)O(\omega^{-1})0, and for sufficiently small O(ω−1)O(\omega^{-1})1 it collapses to a locally exponentially stable equilibrium near the true minimizer. With an added high-pass filter, the admissible adaptation gain depends on a Lipschitz constant but not on the absolute magnitude of the cost, yielding “uniformity in cost amplitude” (Mimmo et al., 2022, Mimmo et al., 2022).

These results motivate a broader interpretation. Exact unbiasedness is the limiting case in which the optimizer itself is the equilibrium of the actual closed loop. Practical unbiasedness is the weaker property that the limiting set can be made arbitrarily small and is not displaced by irrelevant design scales such as cost offsets or finite dither amplitudes. The literature uses both meanings, but the difference is primarily one of asymptotic exactness versus tunable residual error.

4. Objective classes and landscape effects

The strongest recent uES results originally targeted strongly convex maps, but the scope has widened. “Asymptotic, Exponential, and Prescribed-Time Unbiasing in Seeking of Time-Varying Extrema” extends earlier uES designs from strongly convex maps with static optima to a broader class of convex cost functions with time-varying optima diverging at arbitrary rates, even in finite time. The feasibility conditions couple the growth rates of time-varying gains and chirp frequencies to the convexity parameter of the map and to the divergence rate of the optimizer (Yilmaz et al., 2024).

A separate direction asks whether exact derivative estimation is even desirable. In scalar Newton-based ES for strictly but not strongly convex maps, the averaged gradient estimate O(ω−1)O(\omega^{-1})2 is strictly monotone and the averaged Hessian estimate O(ω−1)O(\omega^{-1})3 is strictly positive even when the true Hessian vanishes at the extremum. The resulting perturbation-based NESC is semiglobally practically uniformly asymptotically stable and practically exponentially stable around the averaged equilibrium, while a model-based Newton step would face non-invertibility. The paper explicitly argues that “unbiasedness” of gradient or Hessian estimates is less important than qualitative properties such as monotonicity, positivity, and invertibility (McNamee et al., 2024).

For non-convex maps, the topic splits into two interpretations. The first is global-minimizer seeking under structural assumptions. “Uniform non-convex optimisation via Extremum Seeking” assumes a unique global minimizer and a quasi-convex envelope with bounded local minima depth. The resulting ES dynamics render the global minimizer semi-global practically stable, and with a high-pass filter plus global Lipschitz continuity they yield practical stability with a global domain of attraction (Mimmo et al., 2022). The second is non-local smoothing. “Non-Local Extremum Seeking Based on the Divergence Theorem” constructs a spherical dither that approximates the gradient of the locally averaged objective

O(ω−1)O(\omega^{-1})4

using the identity

O(ω−1)O(\omega^{-1})5

This can eliminate undesired local extrema of O(ω−1)O(\omega^{-1})6, but the optimizer of O(ω−1)O(\omega^{-1})7 need not coincide with that of O(ω−1)O(\omega^{-1})8, so bias is reduced with respect to local traps while potentially introduced with respect to the original objective (Suttner et al., 1 Mar 2026).

Fixed-time and finite-time ES introduce a further distinction. “Fixed-Time Newton-Like Extremum Seeking” gives practical fixed-time convergence to a neighborhood of the optimal point, with a convergence time independent of the initial conditions and the Hessian of the cost function, but not exact asymptotic convergence to the optimizer itself (Poveda et al., 2020). “Multivariable Extremum Seeking Unit-Vector Control Design” similarly proves finite-time stability of the averaged closed-loop error system and practical convergence of the actual system to a neighborhood of the unknown extremum point, with residual error of order O(ω−1)O(\omega^{-1})9 in the state and Jˉa\bar J_a0 in the cost (Silva et al., 9 Apr 2025). In these works, “unbiasedness” is best understood as structural neutrality with respect to curvature or convergence time rather than exact zero steady-state error.

5. Extended settings: delays, PDEs, distributed optimization, and tracking

Unbiased ES has expanded from static scalar maps to complex dynamical and networked settings. In PDE-actuated systems, recent designs compensate delay and diffusion PDE dynamics while ensuring exponential and unbiased convergence to the optimum. The method combines PDE backstepping, exponentially decaying additive dithers, exponentially growing demodulation signals, infinite-dimensional averaging, and local exponential stability of the averaged system. This removes the steady-state displacement that persisted in earlier PDE-compensating ES designs with constant-amplitude dithers (Yilmaz et al., 2024).

Delay-robustness has also reached the discrete-time setting. “Extremum Seeking of Static Maps in the Presence of Unknown Large Time-Varying Delays” presents the first ES algorithm robust with respect to unknown large time-varying delays bounded by known constants, and proves unbiased exponential convergence for Jˉa\bar J_a1-dimensional static quadratic maps. The analysis uses a delay-free transformation and constructive bounds on the controller parameters; larger delay bounds slow the convergence because the admissible small parameter Jˉa\bar J_a2 must be reduced (Jbara et al., 6 Apr 2026).

Distributed optimization introduces an additional layer of bias: local minima of the agents’ objective functions generally do not coincide with the global optimum of the sum cost. “Distributed Time-Varying Optimization via Unbiased Extremum Seeking” therefore combines local-measurement ES with Laplacian-based coordination and scaling transforms of the form

Jˉa\bar J_a3

The resulting continuous-time algorithms operate over weight-balanced, strongly connected directed graphs, use local measurements and neighbor-shared data only, and achieve asymptotic, exponential, or prescribed-time tracking of time-varying optima via chirpy probing and Lie-bracket averaging (Li et al., 26 Sep 2025).

Tracking problems with explicitly time-varying costs fit the same pattern. In a nonlinear chemical-reaction application, a Lie-bracket ES controller minimizes a time-varying tracking cost

Jˉa\bar J_a4

and is shown to keep the state and control within a prescribed neighborhood of the moving optimal curve under suitable assumptions on a steady-state map, local quadratic behavior of the reduced cost, and slow variation of the optimum (Zuyev et al., 10 Jul 2025). More generally, repeated-run ES for unknown dynamic systems can learn controllers that approach the minimal cost for convex objectives and track time-varying LQR-type optimal controllers using only noisy scalar measurements of the performance index (Scheinker et al., 2018).

6. Limitations, ambiguities, and open directions

The topic remains heterogeneous. One recurring ambiguity is definitional: some papers reserve “unbiased” for exact convergence of the actual closed loop to the true optimizer, whereas others use the term for practical convergence with tunable error, for uniformity with respect to cost magnitude or dither amplitude, or for reduction of bias toward local extrema. This suggests that “unbiased ES” is currently a family resemblance rather than a single formal category.

A second limitation is that unbiasing mechanisms introduce their own design burdens. Time-varying gains and chirp frequencies must satisfy nontrivial feasibility conditions tied to convexity, optimizer growth, delay bounds, or PDE eigenvalues (Yilmaz et al., 2024, Jbara et al., 6 Apr 2026). In non-local ES, the smoothing radius Jˉa\bar J_a5 is critical and generally not known online: if Jˉa\bar J_a6 is too small, local extrema survive; if too large, the maximizer of the averaged objective may drift away from that of the original objective (Suttner et al., 1 Mar 2026). In fixed-time Newton-like ES, the result is local because the Riccati inverse-Hessian estimator can have unintended equilibria (Poveda et al., 2020).

Robustness to disturbances, noise, and implementation constraints also remains uneven. Some papers explicitly include bounded measurement errors and derive practical ISS-type estimates for the averaged objective dynamics (Suttner et al., 1 Mar 2026); others note that stochastic robustness, noise-induced bias, actuator bandwidth, and saturation of rapidly increasing gains are still open issues (Yilmaz et al., 2024, Yilmaz et al., 2024). For time-varying tracking, the small parameter Jˉa\bar J_a7 cannot always be taken arbitrarily small because excessively slow adaptation can itself degrade tracking of moving optima (Zuyev et al., 10 Jul 2025).

A plausible implication is that future work will continue to separate three objectives that classical ES treated together: exact asymptotic unbiasing, global attraction in non-convex landscapes, and high-speed transient performance. Recent results show that these objectives can sometimes be combined, but usually only under problem-specific structural assumptions on convexity, smoothness, delay bounds, communication graphs, or averaged objective geometry.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Unbiased Extremum Seeking.