---
title: 'ZOWarmUp: Zeroth-Order Smoothing in ML Training'
url: https://www.emergentmind.com/topics/zowarmup
type: topic
---

# ZOWarmUp: Zeroth-Order Smoothing in ML Training

ZOWarmUp is a term used in recent machine-learning literature for several related stabilization strategies that make zeroth-order or otherwise unstable training procedures practical. In the available arXiv usage, it denotes operating the Zeroth-Order Proximal Operator and Zeroth-Order Proximal Point Algorithm at a fixed positive temperature, a two-stage federated pre-training method that warms up with first-order federated optimization before switching to zeroth-order updates, a learning-rate warmup design for large-scale speech-to-text training, and the sharpness-aware forward-only warm-up stage of SharpZO for CLIP prompt tuning [2605.11929][2509.03503][2505.23420][2506.20990].

## 1. Terminological scope and unifying rationale

Across these usages, ZOWarmUp refers to an initial or persistent smoothing regime introduced to control instability that would otherwise arise from noisy gradient surrogates, heterogeneous hardware, sharp curvature, or deep residual architectures. In the fixed-temperature ZOPO/ZOPPA setting, the relevant mechanism is a positive temperature $\tau$ that stabilizes sampling and reinterprets the method as exact optimization on a smoothed objective. In federated learning, the mechanism is a first-order warm-up on high-resource clients followed by zeroth-order participation of all clients. In large-scale speech-to-text, the mechanism is a deliberately conservative warmup trajectory for the learning rate. In SharpZO, the mechanism is a sharpness-aware CMA-ES phase that precedes sparse local zeroth-order search [2605.11929][2509.03503][2505.23420][2506.20990].

These works do not assign a single universal definition to the term. Rather, they use it for a family of procedures that postpone or regularize the most variance-sensitive part of optimization. This suggests a broader interpretation of ZOWarmUp as a stabilization layer for settings where direct zeroth-order or high-LR training is empirically fragile.

## 2. Fixed-temperature ZOPO/ZOPPA as “ZOWarmUp”

In "Convergence of zeroth-order proximal point algorithms in the high-temperature regime" [2605.11929], ZOWarmUp denotes operating the Zeroth-Order Proximal Operator (ZOPO) and Zeroth-Order Proximal Point Algorithm (ZOPPA) at a fixed positive temperature $\tau$. For $\lambda>0$ and $\tau>0$, the ZOPO is defined by
$$
{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.
$$
Equivalently, it is the posterior mean under the Gibbs/Boltzmann density
$$
\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.
$$
The associated soft Moreau envelope is
$$
f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),
$$
and satisfies
$$
\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).
$$
Hence ZOPPA,
$$
x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),
$$
is exactly gradient descent with step $\lambda$ on $f^{\lambda,\tau}$:
$$
x^{k+1}=x^k-\lambda\nabla f^{\lambda,\tau}(x^k).
$$

A central result is that, at fixed $\tau>0$, ZOPPA is not an inexact proximal method on the original function. It is an exact proximal point method on an auxiliary objective $H_{\lambda,\tau}$ and exact gradient descent on the smoothed objective $f^{\lambda,\tau}$. The paper emphasizes that no inexactness accumulates in this regime. Under the standing assumptions that $f:\mathbb{R}^d\to\mathbb{R}$ is continuous, attains a minimum, and satisfies $e^{-f/\tau}\in L^1(\mathbb{R}^d)$, the iterates converge to a fixed point $x_{\lambda,\tau}^\star$ satisfying
$$
{}^{\tau}_{\lambda,f}(x_{\lambda,\tau}^\star)=x_{\lambda,\tau}^\star
\iff
\nabla f^{\lambda,\tau}(x_{\lambda,\tau}^\star)=0,
$$
with $\|\nabla f^{\lambda,\tau}(x^k)\|\to 0$ and $\sum_k\|\nabla f^{\lambda,\tau}(x^k)\|<\infty$. The sufficient descent identity
$$
f^{\lambda,\tau}(x^{k}) \le f^{\lambda,\tau}(x^{k-1}) - \frac{1}{2\lambda}\|x^{k}-x^{k-1}\|^2
$$
supports the finite-length argument based on analyticity and the Kurdyka-Łojasiewicz property.

The same work derives explicit bias bounds linking stationarity of the smoothed problem back to stationarity of the original objective. When $f$ is $L$-smooth and $\lambda<1/L$,
$$
\|\nabla f(x^k)\| \le (1+\lambda L)\|\nabla f^{\lambda,\tau}(x^k)\| + L\sqrt{\frac{d\,\lambda\,\tau}{1-\lambda L}}.
$$
Related bounds are given for the $G$-Lipschitz case and for quadratic-plus-bounded perturbations $f(y)= (\kappa/2)\|y\|^2 + V(y)$ with finite $\operatorname{osc}(V)$.

A further theme is convexification at fixed temperature. The Hessian identity
$$
\nabla^2 f^{\lambda,\tau}(x)=\frac{1}{\lambda}I-\frac{1}{\lambda^2\tau}\Sigma_{\lambda,\tau}(x),
$$
shows that, as $\lambda$ increases, the negative curvature correction shrinks faster than the positive term. Concrete conditions are stated under which $f^{\lambda,\tau}$ becomes convex on a ball, or globally convex in the quadratic-plus-bounded-$V$ case. In the convex regime, ${}^{\tau}_{\lambda,f}$ is firmly nonexpansive, ZOPPA converges to $x^\star_{\lambda,\tau}\in\arg\min f^{\lambda,\tau}$, and the rates
$$
f^{\lambda,\tau}(x^k)-\min f^{\lambda,\tau}\le \frac{\operatorname{dist}(x^0,X^{\star}_{\lambda,\tau})^2}{2\lambda k},
$$
and
$$
\|\nabla f^{\lambda,\tau}(x^k)\|\le \frac{2\,\operatorname{dist}(x^0,X^{\star}_{\lambda,\tau})}{\lambda(k+1)}
$$
hold.

The sampled method S-ZOPPA uses the self-normalized importance sampling estimator
$$
{}^{\tau,N}_{\lambda,f}(x)=\frac{\sum_{i=1}^N y_i\,\exp(-f(y_i)/\tau)}{\sum_{i=1}^N \exp(-f(y_i)/\tau)},
\qquad
y_i\sim\mathcal{N}(x,\lambda\tau I),
$$
which is biased but consistent and has asymptotic bias and mean-square error of order $N^{-1}$ at fixed $x,\lambda,\tau$. The paper’s practical guidance is to keep $\tau>0$ fixed, start with a relatively large $\lambda$, then decrease $\lambda$ over time while using effective sample size diagnostics to choose $N$. This is the high-temperature “warm-up” interpretation of ZOWarmUp.

## 3. ZOWarmUp in federated pre-training

In "Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients" [2509.03503], ZOWarmUp is a federated, memory-efficient zeroth-order optimizer intended for training from random initialization when many edge devices are below the memory or communication threshold required by standard federated learning. The system contains a server and $K$ clients with non-IID, imbalanced data. Clients are split into high-resource clients $H$, which can run standard federated learning with backpropagation and gradient communication, and low-resource clients $L$, which cannot store backpropagation activations nor exchange full gradients or weights at typical model sizes.

The method has two stages. In step 1, the server runs standard federated training such as FedAvg or FedAdam using only the high-resource set $H$. In step 2, all clients switch to zeroth-order optimization. The paper reports that, empirically, mixing first-order and zeroth-order updates during the second stage degrades performance; best results are obtained when all participating clients use zeroth-order updates in step 2. The pivot point is task- and resource-split dependent; in experiments, 200–300 rounds of warm-up worked well, with the stated trade-off that too little warm-up does not stabilize zeroth-order updates, while too much warm-up withholds data from low-resource clients during a critical learning period.

The zeroth-order estimator is the two-point antithetic estimator
$$
g(w;u)=\frac{f(w+\mu u)-f(w-\mu u)}{2\mu}\cdot u,
$$
and, for $m$ directions,
$$
\hat g(w)=\frac{1}{m}\sum_{i=1}^m \frac{f(w+\mu u_i)-f(w-\mu u_i)}{2\mu}\cdot u_i.
$$
In ZOWarmUp, the perturbation directions are drawn from a scaled Rademacher distribution rather than a Gaussian, which the paper states empirically reduces variance and improves stability. During the zeroth-order stage the server broadcasts a small set of random seeds, clients reconstruct the perturbation directions locally, perform antithetic forward evaluations, and return only scalar function-difference values. Uplink from each client is therefore $O(m)$ scalars rather than $O(d)$-dimensional gradients.

Aggregation is FedAvg-style in the warm-up stage,
$$
w^{(t)} \leftarrow \sum_{k \in P} p_k w_k^{(t)}, \qquad
p_k = \frac{n_k}{\sum_{r\in P} n_r},
$$
where $P\subseteq H$ is the sampled set of high-resource clients. In the zeroth-order stage, each participating client returns $\delta_{j,i}=f_j(w+\mu u_i)-f_j(w-\mu u_i)$, the global scalar is divided by $2\mu$, and the aggregated direction is reconstructed from the shared seeds before the local update
$$
w^{(t+1)} = w^{(t)}-\eta \hat g.
$$
The recommended zeroth-order regime uses single-step local updates per round and full local data, because client drift is exacerbated by noisy zeroth-order gradients if multiple local steps are taken.

The communication and memory analysis is one of the method’s principal claims. For ResNet18, first-order federated learning requires approximately $44.7$ MB uplink and $44.7$ MB downlink per round per client, and about $533.2$ MB on-device memory per client per round. The zeroth-order stage requires approximately $4\cdot S$ bytes uplink per client, approximately $4\cdot S\cdot K$ bytes downlink, and about $89.4$ MB on-device memory. The paper summarizes this as roughly a $6\times$ reduction in on-device memory and orders-of-magnitude communication savings.

Experiments use CIFAR-10 and ImageNet32, 50 clients with Dirichlet $\alpha=0.1$, ResNet18 and ViT-B/16, and resource splits $\{10/90,30/70,50/50,70/30,90/10\}$ for high/low percentages. The standard configuration uses 200 rounds of FedAvg warm-up, followed by 300 zeroth-order rounds with batch size equal to the full local dataset, $S=3$ directions, $\tau=0.75$, and $\mu\approx\epsilon=10^{-4}$. Reported findings include consistent improvement over excluding low-resource clients, a visible accuracy jump at the pivot when low-resource clients join, degradation when increasing the number of local zeroth-order steps, and greater robustness of FedAvg than FedAdam in the zeroth-order stage.

## 4. ZOWarmUp as a learning-rate schedule for large-scale speech-to-text

In "The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence" [2505.23420], ZOWarmUp refers to a learning-rate warmup design for large-scale speech-to-text training with deep Conformer or Branchformer-style encoders. The problem setting is a Conformer encoder with 24 layers, a Transformer decoder with 12 layers, two 1D convolutional layers for $4\times$ subsampling, $d_{\text{model}}=1024$, feed-forward hidden dimension $4096$, 16 attention heads, and approximately $878$M parameters. The training corpus is approximately $150$k hours of English and Italian speech, and each run uses $16\times$ A100 GPUs with 64GB VRAM, $176{,}208$ batches per epoch, and about six days per run.

The baseline inverse-square-root schedule with linear warmup is
$$
\eta(i)=\eta_{\max}\cdot \min\left(\frac{i}{T_w}, \sqrt{\frac{T_w}{i}}\right),
$$
with $T_w=50{,}000$ and $\eta_{\max}=2\times 10^{-4}$. The paper states that models with more than 18 encoder layers diverged under this standard schedule, whereas shallower models converged. The compared alternatives are a piecewise-linear double warmup, a polynomial warmup with $\alpha=1.5$, and an exponential warmup with $\alpha=1.5$. The exponential warmup is
$$
\eta(i)=\eta_{\max}\cdot \frac{\exp\!\left(\alpha\, \frac{i}{T_w}\right)-1}{\exp(\alpha)-1},
\qquad 0\le i\le T_w,
$$
followed by inverse-square-root decay. The double-linear schedule uses $T_1=T_2=25{,}000$ and $\eta_{\mid}=\eta_{\max}/10$.

The convergence evidence is framed in terms of perplexity and gradient norms. Successful schedules keep gradient norms below approximately $25$ after the initial steps, whereas the failing schedules show spikes larger than $100$–$200$ around $25$k–$30$k steps. Exponential warmup converges and reaches earlier “step-like” perplexity drops than the piecewise-linear schedule: approximately $20$k versus $23$k steps for English and approximately $22$k versus $26$k steps for Italian. Polynomial warmup with $\alpha=1.5$ diverges.

Long-run performance is different from early convergence speed. At $170$k steps, piecewise-linear warmup matches or slightly surpasses exponential warmup on word error rate. The reported averages are $13.8$ for piecewise-linear and $14.3$ for exponential, with pairwise values including CommonVoice English $18.4$ versus $19.1$, CommonVoice Italian $13.7$ versus $14.3$, MLS English $7.4$ versus $7.5$, and VoxPopuli English $8.3$ versus $8.6$. The paper’s two principal findings are that large-scale speech-to-text training demands sub-exponential learning-rate warmup for stable convergence, and that a higher learning rate in the warmup phase accelerates initial convergence but does not improve final performance.

The recommended ZOWarmUp designs are therefore either exponential warmup with $\alpha\approx 1.5$ or double-linear warmup with $\eta_{\mid}=\eta_{\max}/10$ and $T_1=T_2=T_w/2$, followed by inverse-square-root decay and gradient-norm monitoring. This usage of the term is not a zeroth-order optimizer in the strict estimator sense; it is a schedule-level stabilization device for deep sequence models.

## 5. ZOWarmUp inside SharpZO for forward-only prompt tuning

In "SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes" [2506.20990], ZOWarmUp is the first stage of SharpZO: a sharpness-aware, forward-only initialization that globally explores and smooths the loss landscape before sparse local zeroth-order optimization. The setting is parameter-efficient CLIP prompt tuning with frozen model weights. The trainable prompt parameter is a low-dimensional vector $w\in\mathbb{R}^d$ with $d=512$, projected into the text embedding space by a fixed random projection $A\in\mathbb{R}^{m\times d}$:
$$
\bar p = wA^\top \in \mathbb{R}^m,
\qquad
p^k=[p_0+\bar p,\; c^k].
$$

The warm-up objective is motivated by sharpness-aware minimization:
$$
\min_w \mathcal{L}^{\text{SAM}_\rho}(w),
\qquad
\mathcal{L}^{\text{SAM}_\rho}(w)=\max_{\|\epsilon\|_2\le \rho}\mathcal{L}(w+\epsilon),
$$
with approximate adversarial perturbation
$$
\epsilon^*\approx \rho \cdot \frac{\nabla \mathcal{L}(w)}{\|\nabla \mathcal{L}(w)\|_2}.
$$
Because backpropagation is unavailable, the gradient is approximated by coordinate-wise two-sided finite differences,
$$
\hat\nabla\mathcal{L}(w)=\sum_{i=1}^d
\left[\frac{\mathcal{L}(w+\mu e_i)-\mathcal{L}(w-\mu e_i)}{2\mu}\right]e_i.
$$
This requires $2d$ forward passes on the minibatch used to compute $\epsilon^*$. The sharpness-aware CMA-ES population is then sampled as
$$
w_t^i \sim \epsilon^* + \theta_t + \delta_t\,\mathcal{N}(0,C_t),
\qquad i=1,\dots,S,
$$
after which $\theta_t$, $\delta_t$, and $C_t$ are updated by standard CMA-ES ranking rules.

Stage 2 is sparse local zeroth-order search using a masked randomized gradient estimator,
$$
\hat\nabla\mathcal{L}(w)=
\frac{\mathcal{L}(w+\mu\cdot \Omega u)-\mathcal{L}(w-\mu\cdot \Omega u)}{2\mu}\,u,
\qquad
u\sim\mathcal{N}(0,I),
$$
and the update
$$
w_{t+1}\leftarrow w_t-\eta \hat\nabla\mathcal{L}(w_t).
$$
The pruning mask is computed from a second-order sensitivity proxy,
$$
\Omega = |w|^2 \cdot z\!\left(\mathbb{E}_{x\sim\mathcal{D}}[\hat\nabla\mathcal{L}(w;x)^2]\right),
\qquad
z(g^2)=\frac{g^2-\mu_g}{\sigma_g}.
$$
The paper reports that, after warm-up, $q=1$ query per step suffices for the local zeroth-order stage.

The theoretical analysis assumes the Polyak-Łojasiewicz condition and Lipschitz smoothness. A per-step bound is stated for the sharpness-aware CMA-ES stage, another for the masked zeroth-order SGD stage, and the main theorem gives a total-step complexity of
$$
t \approx T_c + \mathcal{O}\!\left(\frac{1}{\eta \mu}\,\log\!\left(\frac{L^3\,\eta^2\,\rho^2}{\epsilon}\right)\right).
$$
The interpretation given in the paper is that warm-up reduces the starting gap geometrically and lowers the effective smoothness seen by the zeroth-order phase.

The standard implementation uses population size $S=40$, sharpness radius $\rho=0.1$, intrinsic dimension $d=512$, prompt length $4$, early stopping if validation accuracy fails to improve by more than $0.01$ for 10 consecutive steps, and typical $T_c<100$. On 11 few-shot classification datasets, SharpZO reports average accuracy $75.64\%$ on ViT-B/16 versus $70.57\%$ for ZIP official and $68.20\%$ for BlackVIP, and $69.76\%$ on RN50 versus $64.08\%$ for ZIP and $60.84\%$ for BlackVIP. Time-to-test-accuracy results on a single A100-40G are also reported, with SharpZO substantially faster than ZIP and BlackVIP on ImageNet, Pets, DTD, and EuroSAT.

## 6. Misconceptions, limitations, and cross-domain significance

Several misconceptions are addressed directly by these works. In the fixed-temperature proximal setting, the limit $\tau\to 0$ is not presented as an always-desirable operational regime. Deterministically, it recovers the exact proximal operator, but with sampling the self-normalized importance sampler suffers weight collapse, $\widehat{\mathrm{ESS}}\to 1$, and escape probabilities require exponentially large budgets in dimension [2605.11929]. In federated training, zeroth-order optimization is not portrayed as a drop-in replacement for first-order training from random initialization; the paper states that starting from random initialization with zeroth-order methods often fails to converge, and that warm-up plus variance reduction are what make participation of low-resource clients practical [2509.03503]. In large-scale speech-to-text, simply lowering the peak learning rate is not the recommended solution because it degrades final quality; the evidence instead points to warmup shape at fixed peak learning rate [2505.23420]. In SharpZO, warm-up is beneficial but not free: coordinate-wise gradient estimation for $\epsilon^*$ and mask recomputation costs $2d$ forward passes, and too large a sharpness radius, such as $\rho=0.5$, over-smooths and degrades accuracy [2506.20990].

The limitations are similarly domain-specific. The federated method does not provide new formal convergence guarantees for zeroth-order training from scratch under non-IID heterogeneity, and its performance depends on the hyperparameters $\mu$, $\tau$, and $S$ [2509.03503]. The speech study evaluates only a restricted set of schedules, with polynomial warmup tested only at $\alpha=1.5$, and it leaves open whether additional normalization inside Conformer blocks could mitigate instability [2505.23420]. SharpZO is most naturally suited to low-dimensional prompt tuning rather than full-model optimization because the warm-up cost scales with $d$ [2506.20990]. The fixed-temperature ZOPO/ZOPPA analysis is strongest in the regime where $\tau$ remains positive and sampling remains stable; it explicitly distinguishes this from the computationally unsustainable vanishing-temperature regime [2605.11929].

Taken together, these usages suggest a coherent design pattern: ZOWarmUp denotes a deliberate early-stage or persistent smoothing mechanism that trades aggressive raw optimization for stability, lower estimator variance, broader device participation, or safer traversal of sharp nonconvex landscapes. The common emphasis is not on eliminating approximation, but on making the approximate regime analytically interpretable and computationally dependable.

Source: https://www.emergentmind.com/topics/zowarmup