RESZO: Regression-Based Single-Point ZO
- Regression-Based Single-Point Zeroth-Order Optimization (RESZO) is a derivative-free method that uses regression on historical function evaluations to construct surrogate models and estimate gradients with reduced variance.
- It employs both linear and quadratic surrogate models to capture gradient and curvature information, achieving convergence rates similar to two-point methods while requiring only one function query per iteration.
- RESZO is particularly effective in online, black-box, and simulation-driven scenarios where obtaining multiple function evaluations is impractical or costly.
Regression-Based Single-Point Zeroth-Order Optimization (RESZO) is a class of derivative-free optimization algorithms designed for settings where only a single function evaluation is feasible at each iteration, such as online, black-box, and simulation-driven optimization. The key innovation of RESZO is the use of regression over multiple historical function evaluations to construct local surrogate models, whose gradients serve as low-variance descent directions. This approach achieves convergence rates and query complexities comparable to two-point zeroth-order methods while maintaining the practical and statistical efficiency of single-point evaluations (Chen et al., 6 Jul 2025).
1. Core Principles and Algorithmic Framework
Traditional single-point zeroth-order (SZO) methods estimate gradients using a single sample, e.g., for drawn from a sphere or normal distribution, discarding all previous information. This produces high-variance estimates, leading to slow convergence needing queries to reach stationarity for smooth nonconvex objectives. In contrast, RESZO reuses the most recent function evaluations to fit a local surrogate model by least-squares regression, then takes the surrogate’s gradient as a descent direction. By aggregating historical information, both variance and bias are controlled, accelerating convergence with only one new function evaluation per step.
There are two principal RESZO variants:
- Linear RESZO (L-RESZO): Fits a local linear surrogate around the current perturbed point using recent samples.
- Quadratic RESZO (Q-RESZO): Fits a local quadratic surrogate with a diagonal Hessian to capture basic curvature information.
At each iteration :
- Sample or and set .
- Query .
- Fit a surrogate function 0 using 1 via least-squares regression.
- Update 2.
This regression strategy allows RESZO to leverage the information content of multiple, costly function queries for each update, closing the gap to multi-query (two-point) methods (Chen et al., 6 Jul 2025).
2. Surrogate Model Construction and Algorithmic Implementation
The surrogate at time 3 is built from 4 perturbed points and corresponding function values.
- Linear surrogate:
5 The coefficient 6 (gradient estimate) and offset 7 are given by the least-squares solution:
8
where 9, 0.
- Quadratic surrogate:
Fits a diagonal-Hessian quadratic form,
1
with regression matrices extended accordingly.
Pseudocode for L-RESZO:
3 Q-RESZO follows the same structure with the regression matrix 2 augmented by squared terms for diagonal curvature.
3. Theoretical Guarantees and Convergence Analysis
Under standard smoothness assumptions, the regression-based gradient 3 approximates 4 with error controlled by the window size 5, the step-size 6, and the geometry of the perturbations. Theoretical results for L-RESZO include:
- Gradient-Error Control:
Under 7-smoothness, there exists 8 such that for all 9,
0
for a dimension- and schedule-dependent 1.
- Smooth Nonconvex Case:
For 2 and 3,
4
- Strongly Convex Case:
For smooth 5-strongly convex objectives,
6
- Query Complexity:
| Setting | Two-point ZO | L-RESZO | |------------------------------|--------------|---------------| | Smooth nonconvex | 7 | 8 | | Smooth 9-strongly convex | 0 | 1 |
Empirically, 2 behaves as 3. This suggests that in high dimensions, L-RESZO achieves query complexity comparable (up to a moderate factor) to two-point ZO methods, outperforming standard SZO by a significant margin (Chen et al., 6 Jul 2025).
4. Empirical Performance and Practical Considerations
Comprehensive experiments on noiseless ridge regression, logistic regression, Rosenbrock, and neural network training with 4–5 confirm that both L-RESZO and Q-RESZO converge at essentially the same iteration-rate as two-point ZO, while using only one query per step. Thus, in terms of function query complexity, RESZO is approximately twice as efficient. Both RESZO variants also substantially outperform residual-feedback SZO. Q-RESZO demonstrates slightly faster convergence than L-RESZO due to access to basic curvature information.
Stability and precision are sensitive to the perturbation radius 6:
- 7 causes oscillations or divergence.
- Small, positive 8 increases precision but can hurt stability if too small.
- Adapting 9 provides a balance between stability and optimality.
- Window size: 0 is necessary for full-rank surrogate fitting; 1 is used in practice.
- Overhead: Each iteration requires an 2 or 3 least-squares regression, which can be efficiently updated via rank-one matrix updates.
5. Advantages, Limitations, and Comparison
Advantages
- Single function query per step with far superior variance and convergence properties than classic one-point estimators.
- Systematic reuse of historical data: All available function calls are utilized for each gradient estimation.
- Rates matching two-point ZO: Up to a moderate, empirically mild factor (4).
Limitations
- Assumption A2 dependence: The full theoretical guarantee requires that regression error, as encapsulated by 5, remains bounded—a property empirically observed, but without sharp theoretical bounds for large 6.
- Noiseless analysis: Current convergence results apply only in deterministic function settings.
- Storage and batch-size: Maintaining a buffer of at least 7 past queries is necessary for surrogate regression.
Comparison with other ZO methods
| Method | Queries per step | Uses history | Query complexity (nonconvex) |
|---|---|---|---|
| Classic SZO | 1 | No | 8 |
| Two-point ZO | 2 | Not required | 9 |
| Residual-feedback SZO | 1 | Previous eval only | Improved, but not regression-based |
| RESZO (proposed) | 1 | Yes (window 0) | 1 |
6. Applications and Extensions
RESZO is particularly advantageous in settings where only single function queries are feasible at each iteration:
- Online and dynamic optimization, where the objective may change over time and repeated querying is impossible.
- Bandit settings, expensive simulation, and hyperparameter tuning.
- Reinforcement learning and power systems control, where function evaluation is costly or resource-limited.
- Safety-critical control systems, where repeated, identical actions are not permissible.
A plausible implication is the applicability of RESZO to reinforcement learning and simulation-based policy optimization under severe query limitations.
7. Open Problems and Future Directions
Although RESZO marks a substantial advance for single-point ZO, several technical challenges remain:
- Extending theory to noisy evaluations (stochastic objectives).
- Developing high-probability regret/convergence bounds.
- Rigorous bounding of the regression constant 2 for high-dimensional regimes.
- Improving adaptive strategies for window size and perturbation radius.
- Incorporating variance reduction and acceleration mechanisms.
Potential extensions may include mirror-descent variants, non-Euclidean sampling schemes, or combination with control-oriented feedback designs (Chen et al., 6 Jul 2025).
For the definitive introduction, formal algorithmic details, theoretical analysis, and empirical comparisons, see "Regression-Based Single-Point Zeroth-Order Optimization" (Chen et al., 6 Jul 2025).