Papers
Topics
Authors
Recent
Search
2000 character limit reached

RESZO: Regression-Based Single-Point ZO

Updated 27 November 2025
  • Regression-Based Single-Point Zeroth-Order Optimization (RESZO) is a derivative-free method that uses regression on historical function evaluations to construct surrogate models and estimate gradients with reduced variance.
  • It employs both linear and quadratic surrogate models to capture gradient and curvature information, achieving convergence rates similar to two-point methods while requiring only one function query per iteration.
  • RESZO is particularly effective in online, black-box, and simulation-driven scenarios where obtaining multiple function evaluations is impractical or costly.

Regression-Based Single-Point Zeroth-Order Optimization (RESZO) is a class of derivative-free optimization algorithms designed for settings where only a single function evaluation is feasible at each iteration, such as online, black-box, and simulation-driven optimization. The key innovation of RESZO is the use of regression over multiple historical function evaluations to construct local surrogate models, whose gradients serve as low-variance descent directions. This approach achieves convergence rates and query complexities comparable to two-point zeroth-order methods while maintaining the practical and statistical efficiency of single-point evaluations (Chen et al., 6 Jul 2025).

1. Core Principles and Algorithmic Framework

Traditional single-point zeroth-order (SZO) methods estimate gradients using a single sample, e.g., gt=dδf(xt+δut)utg_t = \frac{d}{\delta} f(x_t + \delta u_t) u_t for utu_t drawn from a sphere or normal distribution, discarding all previous information. This produces high-variance estimates, leading to slow convergence needing O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2}) queries to reach stationarity for smooth nonconvex objectives. In contrast, RESZO reuses the most recent mm function evaluations to fit a local surrogate model by least-squares regression, then takes the surrogate’s gradient as a descent direction. By aggregating historical information, both variance and bias are controlled, accelerating convergence with only one new function evaluation per step.

There are two principal RESZO variants:

  • Linear RESZO (L-RESZO): Fits a local linear surrogate around the current perturbed point using mm recent samples.
  • Quadratic RESZO (Q-RESZO): Fits a local quadratic surrogate with a diagonal Hessian to capture basic curvature information.

At each iteration tt:

  1. Sample ut∼Unif(Sd−1)u_t \sim \text{Unif}(S_{d-1}) or N(0,I)\mathcal{N}(0, I) and set x^t=xt+δut\hat x_t = x_t + \delta u_t.
  2. Query f(x^t)f(\hat x_t).
  3. Fit a surrogate function utu_t0 using utu_t1 via least-squares regression.
  4. Update utu_t2.

This regression strategy allows RESZO to leverage the information content of multiple, costly function queries for each update, closing the gap to multi-query (two-point) methods (Chen et al., 6 Jul 2025).

2. Surrogate Model Construction and Algorithmic Implementation

The surrogate at time utu_t3 is built from utu_t4 perturbed points and corresponding function values.

  • Linear surrogate:

utu_t5 The coefficient utu_t6 (gradient estimate) and offset utu_t7 are given by the least-squares solution:

utu_t8

where utu_t9, O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})0.

  • Quadratic surrogate:

Fits a diagonal-Hessian quadratic form,

O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})1

with regression matrices extended accordingly.

Pseudocode for L-RESZO:

ut∼Unif(Sd−1)u_t \sim \text{Unif}(S_{d-1})3 Q-RESZO follows the same structure with the regression matrix O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})2 augmented by squared terms for diagonal curvature.

3. Theoretical Guarantees and Convergence Analysis

Under standard smoothness assumptions, the regression-based gradient O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})3 approximates O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})4 with error controlled by the window size O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})5, the step-size O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})6, and the geometry of the perturbations. Theoretical results for L-RESZO include:

  • Gradient-Error Control:

Under O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})7-smoothness, there exists O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})8 such that for all O(d3/2/ε3/2)O(d^{3/2}/\varepsilon^{3/2})9,

mm0

for a dimension- and schedule-dependent mm1.

  • Smooth Nonconvex Case:

For mm2 and mm3,

mm4

  • Strongly Convex Case:

For smooth mm5-strongly convex objectives,

mm6

  • Query Complexity:

| Setting | Two-point ZO | L-RESZO | |------------------------------|--------------|---------------| | Smooth nonconvex | mm7 | mm8 | | Smooth mm9-strongly convex | mm0 | mm1 |

Empirically, mm2 behaves as mm3. This suggests that in high dimensions, L-RESZO achieves query complexity comparable (up to a moderate factor) to two-point ZO methods, outperforming standard SZO by a significant margin (Chen et al., 6 Jul 2025).

4. Empirical Performance and Practical Considerations

Comprehensive experiments on noiseless ridge regression, logistic regression, Rosenbrock, and neural network training with mm4–mm5 confirm that both L-RESZO and Q-RESZO converge at essentially the same iteration-rate as two-point ZO, while using only one query per step. Thus, in terms of function query complexity, RESZO is approximately twice as efficient. Both RESZO variants also substantially outperform residual-feedback SZO. Q-RESZO demonstrates slightly faster convergence than L-RESZO due to access to basic curvature information.

Stability and precision are sensitive to the perturbation radius mm6:

  • mm7 causes oscillations or divergence.
  • Small, positive mm8 increases precision but can hurt stability if too small.
  • Adapting mm9 provides a balance between stability and optimality.
  • Window size: tt0 is necessary for full-rank surrogate fitting; tt1 is used in practice.
  • Overhead: Each iteration requires an tt2 or tt3 least-squares regression, which can be efficiently updated via rank-one matrix updates.

5. Advantages, Limitations, and Comparison

Advantages

  • Single function query per step with far superior variance and convergence properties than classic one-point estimators.
  • Systematic reuse of historical data: All available function calls are utilized for each gradient estimation.
  • Rates matching two-point ZO: Up to a moderate, empirically mild factor (tt4).

Limitations

  • Assumption A2 dependence: The full theoretical guarantee requires that regression error, as encapsulated by tt5, remains bounded—a property empirically observed, but without sharp theoretical bounds for large tt6.
  • Noiseless analysis: Current convergence results apply only in deterministic function settings.
  • Storage and batch-size: Maintaining a buffer of at least tt7 past queries is necessary for surrogate regression.

Comparison with other ZO methods

Method Queries per step Uses history Query complexity (nonconvex)
Classic SZO 1 No tt8
Two-point ZO 2 Not required tt9
Residual-feedback SZO 1 Previous eval only Improved, but not regression-based
RESZO (proposed) 1 Yes (window ut∼Unif(Sd−1)u_t \sim \text{Unif}(S_{d-1})0) ut∼Unif(Sd−1)u_t \sim \text{Unif}(S_{d-1})1

6. Applications and Extensions

RESZO is particularly advantageous in settings where only single function queries are feasible at each iteration:

  • Online and dynamic optimization, where the objective may change over time and repeated querying is impossible.
  • Bandit settings, expensive simulation, and hyperparameter tuning.
  • Reinforcement learning and power systems control, where function evaluation is costly or resource-limited.
  • Safety-critical control systems, where repeated, identical actions are not permissible.

A plausible implication is the applicability of RESZO to reinforcement learning and simulation-based policy optimization under severe query limitations.

7. Open Problems and Future Directions

Although RESZO marks a substantial advance for single-point ZO, several technical challenges remain:

  • Extending theory to noisy evaluations (stochastic objectives).
  • Developing high-probability regret/convergence bounds.
  • Rigorous bounding of the regression constant ut∼Unif(Sd−1)u_t \sim \text{Unif}(S_{d-1})2 for high-dimensional regimes.
  • Improving adaptive strategies for window size and perturbation radius.
  • Incorporating variance reduction and acceleration mechanisms.

Potential extensions may include mirror-descent variants, non-Euclidean sampling schemes, or combination with control-oriented feedback designs (Chen et al., 6 Jul 2025).


For the definitive introduction, formal algorithmic details, theoretical analysis, and empirical comparisons, see "Regression-Based Single-Point Zeroth-Order Optimization" (Chen et al., 6 Jul 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Regression-Based Single-Point ZO (RESZO).