Papers
Topics
Authors
Recent
Search
2000 character limit reached

ZOWarmUp: Zeroth-Order Smoothing in ML Training

Updated 10 July 2026
  • The paper shows that ZOWarmUp methods stabilize training by introducing smoothing regimes to control gradient noise and sharp curvature in nonconvex objectives.
  • ZOWarmUp spans various approaches—from fixed-temperature proximal operators and federated warm-up stages to learning-rate schedules—each tailored to mitigate instability.
  • These strategies yield practical benefits, including improved convergence rates, resource efficiency, and robust performance across diverse ML domains.

ZOWarmUp is a term used in recent machine-learning literature for several related stabilization strategies that make zeroth-order or otherwise unstable training procedures practical. In the available arXiv usage, it denotes operating the Zeroth-Order Proximal Operator and Zeroth-Order Proximal Point Algorithm at a fixed positive temperature, a two-stage federated pre-training method that warms up with first-order federated optimization before switching to zeroth-order updates, a learning-rate warmup design for large-scale speech-to-text training, and the sharpness-aware forward-only warm-up stage of SharpZO for CLIP prompt tuning (Naldi et al., 12 May 2026, Legate et al., 3 Sep 2025, Gaido et al., 29 May 2025, Yang et al., 26 Jun 2025).

1. Terminological scope and unifying rationale

Across these usages, ZOWarmUp refers to an initial or persistent smoothing regime introduced to control instability that would otherwise arise from noisy gradient surrogates, heterogeneous hardware, sharp curvature, or deep residual architectures. In the fixed-temperature ZOPO/ZOPPA setting, the relevant mechanism is a positive temperature τ\tau that stabilizes sampling and reinterprets the method as exact optimization on a smoothed objective. In federated learning, the mechanism is a first-order warm-up on high-resource clients followed by zeroth-order participation of all clients. In large-scale speech-to-text, the mechanism is a deliberately conservative warmup trajectory for the learning rate. In SharpZO, the mechanism is a sharpness-aware CMA-ES phase that precedes sparse local zeroth-order search (Naldi et al., 12 May 2026, Legate et al., 3 Sep 2025, Gaido et al., 29 May 2025, Yang et al., 26 Jun 2025).

These works do not assign a single universal definition to the term. Rather, they use it for a family of procedures that postpone or regularize the most variance-sensitive part of optimization. This suggests a broader interpretation of ZOWarmUp as a stabilization layer for settings where direct zeroth-order or high-LR training is empirically fragile.

2. Fixed-temperature ZOPO/ZOPPA as “ZOWarmUp”

In "Convergence of zeroth-order proximal point algorithms in the high-temperature regime" (Naldi et al., 12 May 2026), ZOWarmUp denotes operating the Zeroth-Order Proximal Operator (ZOPO) and Zeroth-Order Proximal Point Algorithm (ZOPPA) at a fixed positive temperature τ\tau. For λ>0\lambda>0 and τ>0\tau>0, the ZOPO is defined by

λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.

Equivalently, it is the posterior mean under the Gibbs/Boltzmann density

πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.

The associated soft Moreau envelope is

fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),

and satisfies

∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).

Hence ZOPPA,

xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),

is exactly gradient descent with step λ\lambda on τ\tau0:

τ\tau1

A central result is that, at fixed τ\tau2, ZOPPA is not an inexact proximal method on the original function. It is an exact proximal point method on an auxiliary objective τ\tau3 and exact gradient descent on the smoothed objective τ\tau4. The paper emphasizes that no inexactness accumulates in this regime. Under the standing assumptions that τ\tau5 is continuous, attains a minimum, and satisfies τ\tau6, the iterates converge to a fixed point τ\tau7 satisfying

τ\tau8

with τ\tau9 and λ>0\lambda>00. The sufficient descent identity

λ>0\lambda>01

supports the finite-length argument based on analyticity and the Kurdyka-Łojasiewicz property.

The same work derives explicit bias bounds linking stationarity of the smoothed problem back to stationarity of the original objective. When λ>0\lambda>02 is λ>0\lambda>03-smooth and λ>0\lambda>04,

λ>0\lambda>05

Related bounds are given for the λ>0\lambda>06-Lipschitz case and for quadratic-plus-bounded perturbations λ>0\lambda>07 with finite λ>0\lambda>08.

A further theme is convexification at fixed temperature. The Hessian identity

λ>0\lambda>09

shows that, as τ>0\tau>00 increases, the negative curvature correction shrinks faster than the positive term. Concrete conditions are stated under which τ>0\tau>01 becomes convex on a ball, or globally convex in the quadratic-plus-bounded-τ>0\tau>02 case. In the convex regime, τ>0\tau>03 is firmly nonexpansive, ZOPPA converges to τ>0\tau>04, and the rates

τ>0\tau>05

and

τ>0\tau>06

hold.

The sampled method S-ZOPPA uses the self-normalized importance sampling estimator

τ>0\tau>07

which is biased but consistent and has asymptotic bias and mean-square error of order τ>0\tau>08 at fixed τ>0\tau>09. The paper’s practical guidance is to keep λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.0 fixed, start with a relatively large λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.1, then decrease λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.2 over time while using effective sample size diagnostics to choose λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.3. This is the high-temperature “warm-up” interpretation of ZOWarmUp.

3. ZOWarmUp in federated pre-training

In "Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients" (Legate et al., 3 Sep 2025), ZOWarmUp is a federated, memory-efficient zeroth-order optimizer intended for training from random initialization when many edge devices are below the memory or communication threshold required by standard federated learning. The system contains a server and λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.4 clients with non-IID, imbalanced data. Clients are split into high-resource clients λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.5, which can run standard federated learning with backpropagation and gradient communication, and low-resource clients λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.6, which cannot store backpropagation activations nor exchange full gradients or weights at typical model sizes.

The method has two stages. In step 1, the server runs standard federated training such as FedAvg or FedAdam using only the high-resource set λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.7. In step 2, all clients switch to zeroth-order optimization. The paper reports that, empirically, mixing first-order and zeroth-order updates during the second stage degrades performance; best results are obtained when all participating clients use zeroth-order updates in step 2. The pivot point is task- and resource-split dependent; in experiments, 200–300 rounds of warm-up worked well, with the stated trade-off that too little warm-up does not stabilize zeroth-order updates, while too much warm-up withholds data from low-resource clients during a critical learning period.

The zeroth-order estimator is the two-point antithetic estimator

λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.8

and, for λ,fτ(x)=Ey∼N(x,λτI)[y exp⁡ ⁣(−f(y)/τ)]Ey∼N(x,λτI)[exp⁡ ⁣(−f(y)/τ)].{}^{\tau}_{\lambda,f}(x)=\frac{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[y\,\exp\!\big(-f(y)/\tau\big)\right]}{\mathbb{E}_{y\sim\mathcal{N}(x,\lambda\tau I)}\left[\exp\!\big(-f(y)/\tau\big)\right]}.9 directions,

πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.0

In ZOWarmUp, the perturbation directions are drawn from a scaled Rademacher distribution rather than a Gaussian, which the paper states empirically reduces variance and improves stability. During the zeroth-order stage the server broadcasts a small set of random seeds, clients reconstruct the perturbation directions locally, perform antithetic forward evaluations, and return only scalar function-difference values. Uplink from each client is therefore πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.1 scalars rather than πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.2-dimensional gradients.

Aggregation is FedAvg-style in the warm-up stage,

πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.3

where πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.4 is the sampled set of high-resource clients. In the zeroth-order stage, each participating client returns πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.5, the global scalar is divided by πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.6, and the aggregated direction is reconstructed from the shared seeds before the local update

πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.7

The recommended zeroth-order regime uses single-step local updates per round and full local data, because client drift is exacerbated by noisy zeroth-order gradients if multiple local steps are taken.

The communication and memory analysis is one of the method’s principal claims. For ResNet18, first-order federated learning requires approximately πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.8 MB uplink and πxλ,τ(y):=exp⁡ ⁣(−f(y)+12λ∥y−x∥2τ)Zλ,τ(x).\pi^{\lambda,\tau}_x(y):=\frac{\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)}{Z_{\lambda,\tau}(x)}.9 MB downlink per round per client, and about fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),0 MB on-device memory per client per round. The zeroth-order stage requires approximately fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),1 bytes uplink per client, approximately fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),2 bytes downlink, and about fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),3 MB on-device memory. The paper summarizes this as roughly a fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),4 reduction in on-device memory and orders-of-magnitude communication savings.

Experiments use CIFAR-10 and ImageNet32, 50 clients with Dirichlet fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),5, ResNet18 and ViT-B/16, and resource splits fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),6 for high/low percentages. The standard configuration uses 200 rounds of FedAvg warm-up, followed by 300 zeroth-order rounds with batch size equal to the full local dataset, fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),7 directions, fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),8, and fλ,τ(x)=− τ log⁡((2πλτ)−d/2∫Rdexp⁡ ⁣(−f(y)+12λ∥y−x∥2τ) dy),f^{\lambda,\tau}(x)= -\,\tau\,\log\left((2\pi\lambda\tau)^{-d/2}\int_{\mathbb{R}^d}\exp\!\Big(-\frac{ f(y) + \frac{1}{2\lambda}\lVert y - x \rVert^2 }{\tau}\Big)\,dy\right),9. Reported findings include consistent improvement over excluding low-resource clients, a visible accuracy jump at the pivot when low-resource clients join, degradation when increasing the number of local zeroth-order steps, and greater robustness of FedAvg than FedAdam in the zeroth-order stage.

4. ZOWarmUp as a learning-rate schedule for large-scale speech-to-text

In "The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence" (Gaido et al., 29 May 2025), ZOWarmUp refers to a learning-rate warmup design for large-scale speech-to-text training with deep Conformer or Branchformer-style encoders. The problem setting is a Conformer encoder with 24 layers, a Transformer decoder with 12 layers, two 1D convolutional layers for ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).0 subsampling, ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).1, feed-forward hidden dimension ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).2, 16 attention heads, and approximately ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).3M parameters. The training corpus is approximately ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).4k hours of English and Italian speech, and each run uses ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).5 A100 GPUs with 64GB VRAM, ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).6 batches per epoch, and about six days per run.

The baseline inverse-square-root schedule with linear warmup is

∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).7

with ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).8 and ∇fλ,τ(x)=1λ(x−λ,fτ(x)).\nabla f^{\lambda,\tau}(x)=\frac{1}{\lambda}\Big(x-{}^{\tau}_{\lambda,f}(x)\Big).9. The paper states that models with more than 18 encoder layers diverged under this standard schedule, whereas shallower models converged. The compared alternatives are a piecewise-linear double warmup, a polynomial warmup with xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),0, and an exponential warmup with xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),1. The exponential warmup is

xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),2

followed by inverse-square-root decay. The double-linear schedule uses xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),3 and xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),4.

The convergence evidence is framed in terms of perplexity and gradient norms. Successful schedules keep gradient norms below approximately xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),5 after the initial steps, whereas the failing schedules show spikes larger than xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),6–xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),7 around xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),8k–xk+1=λ,fτ(xk),x^{k+1} = {}^{\tau}_{\lambda,f}(x^k),9k steps. Exponential warmup converges and reaches earlier “step-like” perplexity drops than the piecewise-linear schedule: approximately λ\lambda0k versus λ\lambda1k steps for English and approximately λ\lambda2k versus λ\lambda3k steps for Italian. Polynomial warmup with λ\lambda4 diverges.

Long-run performance is different from early convergence speed. At λ\lambda5k steps, piecewise-linear warmup matches or slightly surpasses exponential warmup on word error rate. The reported averages are λ\lambda6 for piecewise-linear and λ\lambda7 for exponential, with pairwise values including CommonVoice English λ\lambda8 versus λ\lambda9, CommonVoice Italian τ\tau00 versus τ\tau01, MLS English τ\tau02 versus τ\tau03, and VoxPopuli English τ\tau04 versus τ\tau05. The paper’s two principal findings are that large-scale speech-to-text training demands sub-exponential learning-rate warmup for stable convergence, and that a higher learning rate in the warmup phase accelerates initial convergence but does not improve final performance.

The recommended ZOWarmUp designs are therefore either exponential warmup with τ\tau06 or double-linear warmup with τ\tau07 and τ\tau08, followed by inverse-square-root decay and gradient-norm monitoring. This usage of the term is not a zeroth-order optimizer in the strict estimator sense; it is a schedule-level stabilization device for deep sequence models.

5. ZOWarmUp inside SharpZO for forward-only prompt tuning

In "SharpZO: Hybrid Sharpness-Aware Vision LLM Prompt Tuning via Forward-Only Passes" (Yang et al., 26 Jun 2025), ZOWarmUp is the first stage of SharpZO: a sharpness-aware, forward-only initialization that globally explores and smooths the loss landscape before sparse local zeroth-order optimization. The setting is parameter-efficient CLIP prompt tuning with frozen model weights. The trainable prompt parameter is a low-dimensional vector τ\tau09 with τ\tau10, projected into the text embedding space by a fixed random projection τ\tau11:

τ\tau12

The warm-up objective is motivated by sharpness-aware minimization:

τ\tau13

with approximate adversarial perturbation

τ\tau14

Because backpropagation is unavailable, the gradient is approximated by coordinate-wise two-sided finite differences,

τ\tau15

This requires τ\tau16 forward passes on the minibatch used to compute τ\tau17. The sharpness-aware CMA-ES population is then sampled as

τ\tau18

after which τ\tau19, τ\tau20, and τ\tau21 are updated by standard CMA-ES ranking rules.

Stage 2 is sparse local zeroth-order search using a masked randomized gradient estimator,

τ\tau22

and the update

τ\tau23

The pruning mask is computed from a second-order sensitivity proxy,

τ\tau24

The paper reports that, after warm-up, τ\tau25 query per step suffices for the local zeroth-order stage.

The theoretical analysis assumes the Polyak-Łojasiewicz condition and Lipschitz smoothness. A per-step bound is stated for the sharpness-aware CMA-ES stage, another for the masked zeroth-order SGD stage, and the main theorem gives a total-step complexity of

τ\tau26

The interpretation given in the paper is that warm-up reduces the starting gap geometrically and lowers the effective smoothness seen by the zeroth-order phase.

The standard implementation uses population size τ\tau27, sharpness radius τ\tau28, intrinsic dimension τ\tau29, prompt length τ\tau30, early stopping if validation accuracy fails to improve by more than τ\tau31 for 10 consecutive steps, and typical τ\tau32. On 11 few-shot classification datasets, SharpZO reports average accuracy τ\tau33 on ViT-B/16 versus τ\tau34 for ZIP official and τ\tau35 for BlackVIP, and τ\tau36 on RN50 versus τ\tau37 for ZIP and τ\tau38 for BlackVIP. Time-to-test-accuracy results on a single A100-40G are also reported, with SharpZO substantially faster than ZIP and BlackVIP on ImageNet, Pets, DTD, and EuroSAT.

6. Misconceptions, limitations, and cross-domain significance

Several misconceptions are addressed directly by these works. In the fixed-temperature proximal setting, the limit τ\tau39 is not presented as an always-desirable operational regime. Deterministically, it recovers the exact proximal operator, but with sampling the self-normalized importance sampler suffers weight collapse, τ\tau40, and escape probabilities require exponentially large budgets in dimension (Naldi et al., 12 May 2026). In federated training, zeroth-order optimization is not portrayed as a drop-in replacement for first-order training from random initialization; the paper states that starting from random initialization with zeroth-order methods often fails to converge, and that warm-up plus variance reduction are what make participation of low-resource clients practical (Legate et al., 3 Sep 2025). In large-scale speech-to-text, simply lowering the peak learning rate is not the recommended solution because it degrades final quality; the evidence instead points to warmup shape at fixed peak learning rate (Gaido et al., 29 May 2025). In SharpZO, warm-up is beneficial but not free: coordinate-wise gradient estimation for τ\tau41 and mask recomputation costs τ\tau42 forward passes, and too large a sharpness radius, such as τ\tau43, over-smooths and degrades accuracy (Yang et al., 26 Jun 2025).

The limitations are similarly domain-specific. The federated method does not provide new formal convergence guarantees for zeroth-order training from scratch under non-IID heterogeneity, and its performance depends on the hyperparameters τ\tau44, τ\tau45, and τ\tau46 (Legate et al., 3 Sep 2025). The speech study evaluates only a restricted set of schedules, with polynomial warmup tested only at τ\tau47, and it leaves open whether additional normalization inside Conformer blocks could mitigate instability (Gaido et al., 29 May 2025). SharpZO is most naturally suited to low-dimensional prompt tuning rather than full-model optimization because the warm-up cost scales with τ\tau48 (Yang et al., 26 Jun 2025). The fixed-temperature ZOPO/ZOPPA analysis is strongest in the regime where τ\tau49 remains positive and sampling remains stable; it explicitly distinguishes this from the computationally unsustainable vanishing-temperature regime (Naldi et al., 12 May 2026).

Taken together, these usages suggest a coherent design pattern: ZOWarmUp denotes a deliberate early-stage or persistent smoothing mechanism that trades aggressive raw optimization for stability, lower estimator variance, broader device participation, or safer traversal of sharp nonconvex landscapes. The common emphasis is not on eliminating approximation, but on making the approximate regime analytically interpretable and computationally dependable.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ZOWarmUp.