Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning-Based Stackelberg Equilibrium Seeking with Application to Demand-Side Energy Management

Published 1 May 2026 in math.OC | (2605.00588v1)

Abstract: Demand-side management (DSM) enables distribution system operators (DSOs) to steer electricity consumption through dynamic price signals or incentive mechanisms, thereby leveraging end-users' flexibility potential for delivering grid services. The resulting hierarchical interaction between the DSO and the end-users can be formulated as a Stackelberg game, where the operator dynamically sets the prices and the end-users optimally respond to them. Efficiently designing these price signals is challenging, as the users' response models are unknown or difficult to estimate. In this paper, we propose a learning-based zeroth-order algorithm for incentive design, in which the iterative update of the incentive signals is efficiently assisted by a data-driven online estimation of the users' responses. The proposed method is then proven to converge to an equilibrium tariff while allowing the DSO to estimate the decision-making problems at the user level. Moreover, the method preserves users' privacy, as the update rule of the DSO is solely based on observations of communicated end-user actions. Numerical simulations employing real-world data illustrate the efficient convergence of our learning-based proposed method, while significantly reducing the number of required interactions between the DSO and the end-users with respect to the state-of-the-art approach.

Summary

  • The paper introduces a hybrid Stackelberg learning method that combines zeroth-order tariff updates with data-driven estimation of prosumers’ nonsmooth response maps, retaining privacy while providing convergence guarantees to approximate stationary points.
  • Simulations with 10 prosumers, real consumption data, and Belgian day-ahead prices reduce normalized objective error below 10^-4 in about 200 iterations and cut DSO–user interactions by several orders of magnitude versus unassisted zeroth-order optimization.
  • The method depends on local surrogate accuracy and trust-region tuning: limited private-constraint knowledge slows convergence to roughly 700 iterations and 10^-3 accuracy, while overly long surrogate updates can cause divergence.

Overview

This paper addresses the design of dynamic network tariffs in demand-side management (DSM), formulated as a Stackelberg game between a distribution system operator (DSO), acting as leader, and an energy community (EC) of NN prosumers acting as followers. The central obstacle is that the DSO does not know the followers' decision-making model, and existing remedies either require strong information (bilevel/MPEC formulations, Jacobian communication), or preserve privacy but converge slowly (zeroth-order probing). The authors propose a hybrid scheme that alternates zeroth-order tariff updates with data-driven estimation of the parametric lower-level response map, and they support the method with convergence guarantees and simulations on real-world consumption data (2605.00588).

Problem formulation

The followers' game is a generalized Nash game with coupling constraints: each prosumer ii schedules the consumption of mim_i assets over T=24T=24 hourly slots to minimize a weighted combination of discomfort, xiri2\|x_i - r_i\|^2, and energy cost, (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i, subject to local constraints Xi\mathcal{X}_i and the shared capacity constraint iAixib\sum_i A_i x_i \le b. The weighting matrix Πi\Pi_i encodes heterogeneous price sensitivities and contractual rules. As solution concept the authors adopt the variational GNE (vv-GNE), and Lemma 1 establishes that the ii0-GNE set is a singleton given by the projection map

ii1

where ii2 is a polyhedron capturing shared and local constraints. The DSO (leader) selects the tariff ii3 to maximize revenue ii4 while penalizing deviation from a prescribed average tariff and tariff volatility. By Lemma 1, the bilevel Stackelberg problem collapses to the unconstrained nonsmooth nonconvex problem ii5, whose global minima are Stackelberg equilibria; the paper instead targets stationary points, consistent with prior work on follower-agnostic Stackelberg learning.

Methodology

The proposed procedure cycles through three steps: (a) zeroth-order tariff updates with dataset collection, (b) parameter estimation of the lower-level model, and (c) solution of the estimated bilevel problem to produce a warm start for the next round.

Nonsmooth zeroth-order method. The projection in ii6 makes ii7 nonsmooth, violating the smoothness assumption of the zeroth-order method of Maheshwari et al. The authors replace the projection with a smooth squared-softplus penalty of parameter ii8, yielding a smoothed objective ii9 whose stationary points converge to Clarke stationary points as mim_i0. The DSO probes the EC at mim_i1 and a perturbed point mim_i2, forms the two-point estimator mim_i3 from observed (nonsmooth) responses rather than the smoothed ones, and updates mim_i4. Theorem 1 provides the explicit bound

mim_i5

with mim_i6 under step sizes mim_i7 and radii mim_i8. The terms mim_i9 vanish with T=24T=240; notably, T=24T=241 grows with T=24T=242 because the estimator operates on the true nonsmooth responses, but it can be made arbitrarily small by shrinking T=24T=243. This bias term is the main technical departure from the prior method.

Surrogate model estimation and trust region. After T=24T=244 iterations, the DSO fits T=24T=245 by minimizing the empirical suboptimality loss between observed and predicted leader objectives, subject to the parametric projection structure. Theorem 2 establishes only local validity: around each data point T=24T=246 there is a ball of radius T=24T=247 on which the surrogate is T=24T=248-accurate, and consequently the estimated-model optimum T=24T=249 satisfies

xiri2\|x_i - r_i\|^20

This bound is deliberately modest: it guarantees that accepting a candidate point from the surrogate deteriorates the true objective by at most xiri2\|x_i - r_i\|^21, but it does not guarantee improvement. The authors are explicit that the guarantee is local, which motivates the trust-region logic below.

DCP-based surrogate optimization. The estimated bilevel problem admits a difference-of-convex-programming (DCP) reformulation (Proposition 1), solved by successive linearization of the concave part with monotone decrease guarantees. Because the surrogate is reliable only within the trust region, the DCP horizon xiri2\|x_i - r_i\|^22 must be tuned so iterates remain inside xiri2\|x_i - r_i\|^23; since this region is not explicitly characterizable, the authors propose the practical proxy of stopping at the largest xiri2\|x_i - r_i\|^24 with monotone decrease of the true objective — which itself requires querying the users, so xiri2\|x_i - r_i\|^25 cannot be determined a priori. A remark further derives an adaptive rule for choosing the next round length xiri2\|x_i - r_i\|^26 by comparing convergence bounds with and without accepting the candidate xiri2\|x_i - r_i\|^27; the authors note this rule is conservative.

Numerical results

Simulations use the setup of the authors' earlier two-part pricing work: xiri2\|x_i - r_i\|^28 prosumers, each equipped with a shiftable load, a battery, an EV, and rooftop PV, with real consumption data and day-ahead Belgian prices. The leader weights are xiri2\|x_i - r_i\|^29, (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i0, and the optimum (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i1 is computed with Ipopt.

Three findings stand out. First, with the true feasible set (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i2 and optimally tuned (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i3, the hybrid method drives the normalized objective error below (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i4 in roughly 200 iterations under all round lengths (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i5; (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i6 converges fastest (about 60 iterations) because estimation jumps occur more frequently, while the adaptive-(p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i7 rule converges more slowly due to bound conservativeness, though it suppresses the oscillations that fixed (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i8 exhibits near convergence. Second, the comparison against unassisted zeroth-order iterations (bottom of the first figure) shows a reduction in required DSO–user interactions of several orders of magnitude — the paper's strongest empirical claim. Third, relaxing knowledge of the private constraint parameters (p+y)ΠiAixi(p+y)^\top \Pi_i A_i x_i9 (estimated from data as Xi\mathcal{X}_i0) degrades performance: convergence slows to roughly 700 iterations and accuracy deteriorates to Xi\mathcal{X}_i1, still outperforming unassisted zeroth-order optimization. Finally, the Xi\mathcal{X}_i2 study shows the trust-region mechanism is essential: with Xi\mathcal{X}_i3 the surrogate eventually produces upward jumps and divergence (accuracy stalls above Xi\mathcal{X}_i4), while Xi\mathcal{X}_i5 guarantees convergence at Xi\mathcal{X}_i6 accuracy at the cost of speed.

Limitations and open questions

The paper concedes several restrictions at the points where they bear on the results. The convergence guarantee of Theorem 1 applies to the smoothed objective and carries the Xi\mathcal{X}_i7 bias term, so the guarantee is for approximate, not exact, Clarke stationarity. Theorem 2's surrogate validity is local and does not certify improvement of the candidate point, only bounded deterioration; the trust region Xi\mathcal{X}_i8 is not explicitly characterizable, and both Xi\mathcal{X}_i9 and the acceptance test require additional probing of the users, partially eroding the claimed reduction in interactions. The estimation step assumes the parametric projection structure of the response map is known to the DSO, an assumption that holds only approximately in practice. The adaptive rule for iAixib\sum_i A_i x_i \le b0 is empirically conservative and slower than fixed-iAixib\sum_i A_i x_i \le b1 operation. Open questions include a principled, non-conservative criterion for iAixib\sum_i A_i x_i \le b2 and for accepting surrogate candidates without extra queries, and convergence analysis under inexact or noisy estimation of iAixib\sum_i A_i x_i \le b3.

Conclusion

The paper contributes a hybrid framework that augments privacy-preserving zeroth-order Stackelberg learning with online estimation of the followers' parametric response model, extending the underlying method to nonsmooth projection-based reactions with an explicit convergence bound. Simulations on real-world data indicate that the learned surrogate accelerates convergence by several orders of magnitude relative to pure zeroth-order methods, at the cost of local-only surrogate guarantees and empirically tuned trust-region parameters.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.