ZOWarmUp: Zeroth-Order Smoothing in ML Training
- The paper shows that ZOWarmUp methods stabilize training by introducing smoothing regimes to control gradient noise and sharp curvature in nonconvex objectives.
- ZOWarmUp spans various approaches—from fixed-temperature proximal operators and federated warm-up stages to learning-rate schedules—each tailored to mitigate instability.
- These strategies yield practical benefits, including improved convergence rates, resource efficiency, and robust performance across diverse ML domains.
ZOWarmUp is a term used in recent machine-learning literature for several related stabilization strategies that make zeroth-order or otherwise unstable training procedures practical. In the available arXiv usage, it denotes operating the Zeroth-Order Proximal Operator and Zeroth-Order Proximal Point Algorithm at a fixed positive temperature, a two-stage federated pre-training method that warms up with first-order federated optimization before switching to zeroth-order updates, a learning-rate warmup design for large-scale speech-to-text training, and the sharpness-aware forward-only warm-up stage of SharpZO for CLIP prompt tuning (Naldi et al., 12 May 2026, Legate et al., 3 Sep 2025, Gaido et al., 29 May 2025, Yang et al., 26 Jun 2025).
1. Terminological scope and unifying rationale
Across these usages, ZOWarmUp refers to an initial or persistent smoothing regime introduced to control instability that would otherwise arise from noisy gradient surrogates, heterogeneous hardware, sharp curvature, or deep residual architectures. In the fixed-temperature ZOPO/ZOPPA setting, the relevant mechanism is a positive temperature that stabilizes sampling and reinterprets the method as exact optimization on a smoothed objective. In federated learning, the mechanism is a first-order warm-up on high-resource clients followed by zeroth-order participation of all clients. In large-scale speech-to-text, the mechanism is a deliberately conservative warmup trajectory for the learning rate. In SharpZO, the mechanism is a sharpness-aware CMA-ES phase that precedes sparse local zeroth-order search (Naldi et al., 12 May 2026, Legate et al., 3 Sep 2025, Gaido et al., 29 May 2025, Yang et al., 26 Jun 2025).
These works do not assign a single universal definition to the term. Rather, they use it for a family of procedures that postpone or regularize the most variance-sensitive part of optimization. This suggests a broader interpretation of ZOWarmUp as a stabilization layer for settings where direct zeroth-order or high-LR training is empirically fragile.
2. Fixed-temperature ZOPO/ZOPPA as “ZOWarmUp”
In "Convergence of zeroth-order proximal point algorithms in the high-temperature regime" (Naldi et al., 12 May 2026), ZOWarmUp denotes operating the Zeroth-Order Proximal Operator (ZOPO) and Zeroth-Order Proximal Point Algorithm (ZOPPA) at a fixed positive temperature . For and , the ZOPO is defined by
Equivalently, it is the posterior mean under the Gibbs/Boltzmann density
The associated soft Moreau envelope is
and satisfies
Hence ZOPPA,
is exactly gradient descent with step on 0:
1
A central result is that, at fixed 2, ZOPPA is not an inexact proximal method on the original function. It is an exact proximal point method on an auxiliary objective 3 and exact gradient descent on the smoothed objective 4. The paper emphasizes that no inexactness accumulates in this regime. Under the standing assumptions that 5 is continuous, attains a minimum, and satisfies 6, the iterates converge to a fixed point 7 satisfying
8
with 9 and 0. The sufficient descent identity
1
supports the finite-length argument based on analyticity and the Kurdyka-Łojasiewicz property.
The same work derives explicit bias bounds linking stationarity of the smoothed problem back to stationarity of the original objective. When 2 is 3-smooth and 4,
5
Related bounds are given for the 6-Lipschitz case and for quadratic-plus-bounded perturbations 7 with finite 8.
A further theme is convexification at fixed temperature. The Hessian identity
9
shows that, as 0 increases, the negative curvature correction shrinks faster than the positive term. Concrete conditions are stated under which 1 becomes convex on a ball, or globally convex in the quadratic-plus-bounded-2 case. In the convex regime, 3 is firmly nonexpansive, ZOPPA converges to 4, and the rates
5
and
6
hold.
The sampled method S-ZOPPA uses the self-normalized importance sampling estimator
7
which is biased but consistent and has asymptotic bias and mean-square error of order 8 at fixed 9. The paper’s practical guidance is to keep 0 fixed, start with a relatively large 1, then decrease 2 over time while using effective sample size diagnostics to choose 3. This is the high-temperature “warm-up” interpretation of ZOWarmUp.
3. ZOWarmUp in federated pre-training
In "Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients" (Legate et al., 3 Sep 2025), ZOWarmUp is a federated, memory-efficient zeroth-order optimizer intended for training from random initialization when many edge devices are below the memory or communication threshold required by standard federated learning. The system contains a server and 4 clients with non-IID, imbalanced data. Clients are split into high-resource clients 5, which can run standard federated learning with backpropagation and gradient communication, and low-resource clients 6, which cannot store backpropagation activations nor exchange full gradients or weights at typical model sizes.
The method has two stages. In step 1, the server runs standard federated training such as FedAvg or FedAdam using only the high-resource set 7. In step 2, all clients switch to zeroth-order optimization. The paper reports that, empirically, mixing first-order and zeroth-order updates during the second stage degrades performance; best results are obtained when all participating clients use zeroth-order updates in step 2. The pivot point is task- and resource-split dependent; in experiments, 200–300 rounds of warm-up worked well, with the stated trade-off that too little warm-up does not stabilize zeroth-order updates, while too much warm-up withholds data from low-resource clients during a critical learning period.
The zeroth-order estimator is the two-point antithetic estimator
8
and, for 9 directions,
0
In ZOWarmUp, the perturbation directions are drawn from a scaled Rademacher distribution rather than a Gaussian, which the paper states empirically reduces variance and improves stability. During the zeroth-order stage the server broadcasts a small set of random seeds, clients reconstruct the perturbation directions locally, perform antithetic forward evaluations, and return only scalar function-difference values. Uplink from each client is therefore 1 scalars rather than 2-dimensional gradients.
Aggregation is FedAvg-style in the warm-up stage,
3
where 4 is the sampled set of high-resource clients. In the zeroth-order stage, each participating client returns 5, the global scalar is divided by 6, and the aggregated direction is reconstructed from the shared seeds before the local update
7
The recommended zeroth-order regime uses single-step local updates per round and full local data, because client drift is exacerbated by noisy zeroth-order gradients if multiple local steps are taken.
The communication and memory analysis is one of the method’s principal claims. For ResNet18, first-order federated learning requires approximately 8 MB uplink and 9 MB downlink per round per client, and about 0 MB on-device memory per client per round. The zeroth-order stage requires approximately 1 bytes uplink per client, approximately 2 bytes downlink, and about 3 MB on-device memory. The paper summarizes this as roughly a 4 reduction in on-device memory and orders-of-magnitude communication savings.
Experiments use CIFAR-10 and ImageNet32, 50 clients with Dirichlet 5, ResNet18 and ViT-B/16, and resource splits 6 for high/low percentages. The standard configuration uses 200 rounds of FedAvg warm-up, followed by 300 zeroth-order rounds with batch size equal to the full local dataset, 7 directions, 8, and 9. Reported findings include consistent improvement over excluding low-resource clients, a visible accuracy jump at the pivot when low-resource clients join, degradation when increasing the number of local zeroth-order steps, and greater robustness of FedAvg than FedAdam in the zeroth-order stage.
4. ZOWarmUp as a learning-rate schedule for large-scale speech-to-text
In "The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence" (Gaido et al., 29 May 2025), ZOWarmUp refers to a learning-rate warmup design for large-scale speech-to-text training with deep Conformer or Branchformer-style encoders. The problem setting is a Conformer encoder with 24 layers, a Transformer decoder with 12 layers, two 1D convolutional layers for 0 subsampling, 1, feed-forward hidden dimension 2, 16 attention heads, and approximately 3M parameters. The training corpus is approximately 4k hours of English and Italian speech, and each run uses 5 A100 GPUs with 64GB VRAM, 6 batches per epoch, and about six days per run.
The baseline inverse-square-root schedule with linear warmup is
7
with 8 and 9. The paper states that models with more than 18 encoder layers diverged under this standard schedule, whereas shallower models converged. The compared alternatives are a piecewise-linear double warmup, a polynomial warmup with 0, and an exponential warmup with 1. The exponential warmup is
2
followed by inverse-square-root decay. The double-linear schedule uses 3 and 4.
The convergence evidence is framed in terms of perplexity and gradient norms. Successful schedules keep gradient norms below approximately 5 after the initial steps, whereas the failing schedules show spikes larger than 6–7 around 8k–9k steps. Exponential warmup converges and reaches earlier “step-like” perplexity drops than the piecewise-linear schedule: approximately 0k versus 1k steps for English and approximately 2k versus 3k steps for Italian. Polynomial warmup with 4 diverges.
Long-run performance is different from early convergence speed. At 5k steps, piecewise-linear warmup matches or slightly surpasses exponential warmup on word error rate. The reported averages are 6 for piecewise-linear and 7 for exponential, with pairwise values including CommonVoice English 8 versus 9, CommonVoice Italian 00 versus 01, MLS English 02 versus 03, and VoxPopuli English 04 versus 05. The paper’s two principal findings are that large-scale speech-to-text training demands sub-exponential learning-rate warmup for stable convergence, and that a higher learning rate in the warmup phase accelerates initial convergence but does not improve final performance.
The recommended ZOWarmUp designs are therefore either exponential warmup with 06 or double-linear warmup with 07 and 08, followed by inverse-square-root decay and gradient-norm monitoring. This usage of the term is not a zeroth-order optimizer in the strict estimator sense; it is a schedule-level stabilization device for deep sequence models.
5. ZOWarmUp inside SharpZO for forward-only prompt tuning
In "SharpZO: Hybrid Sharpness-Aware Vision LLM Prompt Tuning via Forward-Only Passes" (Yang et al., 26 Jun 2025), ZOWarmUp is the first stage of SharpZO: a sharpness-aware, forward-only initialization that globally explores and smooths the loss landscape before sparse local zeroth-order optimization. The setting is parameter-efficient CLIP prompt tuning with frozen model weights. The trainable prompt parameter is a low-dimensional vector 09 with 10, projected into the text embedding space by a fixed random projection 11:
12
The warm-up objective is motivated by sharpness-aware minimization:
13
with approximate adversarial perturbation
14
Because backpropagation is unavailable, the gradient is approximated by coordinate-wise two-sided finite differences,
15
This requires 16 forward passes on the minibatch used to compute 17. The sharpness-aware CMA-ES population is then sampled as
18
after which 19, 20, and 21 are updated by standard CMA-ES ranking rules.
Stage 2 is sparse local zeroth-order search using a masked randomized gradient estimator,
22
and the update
23
The pruning mask is computed from a second-order sensitivity proxy,
24
The paper reports that, after warm-up, 25 query per step suffices for the local zeroth-order stage.
The theoretical analysis assumes the Polyak-Łojasiewicz condition and Lipschitz smoothness. A per-step bound is stated for the sharpness-aware CMA-ES stage, another for the masked zeroth-order SGD stage, and the main theorem gives a total-step complexity of
26
The interpretation given in the paper is that warm-up reduces the starting gap geometrically and lowers the effective smoothness seen by the zeroth-order phase.
The standard implementation uses population size 27, sharpness radius 28, intrinsic dimension 29, prompt length 30, early stopping if validation accuracy fails to improve by more than 31 for 10 consecutive steps, and typical 32. On 11 few-shot classification datasets, SharpZO reports average accuracy 33 on ViT-B/16 versus 34 for ZIP official and 35 for BlackVIP, and 36 on RN50 versus 37 for ZIP and 38 for BlackVIP. Time-to-test-accuracy results on a single A100-40G are also reported, with SharpZO substantially faster than ZIP and BlackVIP on ImageNet, Pets, DTD, and EuroSAT.
6. Misconceptions, limitations, and cross-domain significance
Several misconceptions are addressed directly by these works. In the fixed-temperature proximal setting, the limit 39 is not presented as an always-desirable operational regime. Deterministically, it recovers the exact proximal operator, but with sampling the self-normalized importance sampler suffers weight collapse, 40, and escape probabilities require exponentially large budgets in dimension (Naldi et al., 12 May 2026). In federated training, zeroth-order optimization is not portrayed as a drop-in replacement for first-order training from random initialization; the paper states that starting from random initialization with zeroth-order methods often fails to converge, and that warm-up plus variance reduction are what make participation of low-resource clients practical (Legate et al., 3 Sep 2025). In large-scale speech-to-text, simply lowering the peak learning rate is not the recommended solution because it degrades final quality; the evidence instead points to warmup shape at fixed peak learning rate (Gaido et al., 29 May 2025). In SharpZO, warm-up is beneficial but not free: coordinate-wise gradient estimation for 41 and mask recomputation costs 42 forward passes, and too large a sharpness radius, such as 43, over-smooths and degrades accuracy (Yang et al., 26 Jun 2025).
The limitations are similarly domain-specific. The federated method does not provide new formal convergence guarantees for zeroth-order training from scratch under non-IID heterogeneity, and its performance depends on the hyperparameters 44, 45, and 46 (Legate et al., 3 Sep 2025). The speech study evaluates only a restricted set of schedules, with polynomial warmup tested only at 47, and it leaves open whether additional normalization inside Conformer blocks could mitigate instability (Gaido et al., 29 May 2025). SharpZO is most naturally suited to low-dimensional prompt tuning rather than full-model optimization because the warm-up cost scales with 48 (Yang et al., 26 Jun 2025). The fixed-temperature ZOPO/ZOPPA analysis is strongest in the regime where 49 remains positive and sampling remains stable; it explicitly distinguishes this from the computationally unsustainable vanishing-temperature regime (Naldi et al., 12 May 2026).
Taken together, these usages suggest a coherent design pattern: ZOWarmUp denotes a deliberate early-stage or persistent smoothing mechanism that trades aggressive raw optimization for stability, lower estimator variance, broader device participation, or safer traversal of sharp nonconvex landscapes. The common emphasis is not on eliminating approximation, but on making the approximate regime analytically interpretable and computationally dependable.