Papers
Topics
Authors
Recent
Search
2000 character limit reached

1.5-SPSA Optimization Method

Updated 6 October 2026
  • 1.5-SPSA is a zero-order optimization method specifically designed for neural-network training, characterized by adding a central clean evaluation per optimization step to estimate and use directional curvature.
  • 1.5-SPSA achieves improved performance by reweighting gradient directions based on curvature estimates, which enhances the efficiency of neural network training, particularly in memory-constrained environments. It requires fewer forward passes per step but may be sensitive to hyperparameter settings.
  • This method parallelizes probe evaluations across GPUs to optimize training time, effectively treating positive and negative curvature magnitudes equally, and resulting in moderate suppression of extreme curvatures, which retains inference-mode retention characteristics in training.

1.5-SPSA is a zero-order, inference-mode optimization method introduced for large-scale neural-network training. It extends ordinary one-probe SPSA (1SPSA) with one additional clean, unperturbed forward evaluation per optimization step. This shared center-point evaluation enables estimation of directional curvature for each random perturbation, which is then used to reweight probe contributions. The method does not construct or store a parameter-space Hessian; its preconditioner operates in the space of random perturbation directions (Chaubard et al., 29 Sep 2026).

1. Terminology and conceptual position

The name “1.5-SPSA” denotes an intermediate construction between first-order SPSA and a full second-order method. Ordinary two-sided SPSA estimates directional slopes using evaluations at θ+ϵzi\theta+\epsilon z_i and θ−ϵzi\theta-\epsilon z_i. 1.5-SPSA adds the evaluation L(θ)L(\theta), producing the three-point stencil

c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.

For twice-differentiable LL, this quantity approximates the curvature along the probe direction:

c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.

The method is therefore “one-and-a-half” in the sense that it retains SPSA’s first-order random-probe structure while adding a limited second-order correction. It is not a Newton method, does not estimate a full Hessian, and does not compute a diagonal Hessian in parameter coordinates.

The method should be distinguished from several unrelated uses of “1.5-SPSA.” The SPS upgrade paper does not define the term, and the QAOA, beamforming, SPSA-theory, and VQE papers likewise do not introduce a method with this name. In particular, the $3/2$ exponent appearing in finite-shot measurement-complexity analyses is not a fractional-order SPSA algorithm. Guided-SPSA combines parameter-shift gradients and SPSA but is a sample-partitioned hybrid rather than 1.5-SPSA. The method defined here is specifically probe-space preconditioning for zero-order training (Goddard et al., 2014, Chen et al., 2021, Periyasamy et al., 2024, Qin, 10 Aug 2026).

2. Motivation and optimization setting

Backpropagation requires storing intermediate activations and, when used with Adam, maintaining first- and second-moment optimizer states. The paper gives training OPT-30B with Adam, batch size 8, and sequence length 2048 as requiring approximately 600 GB of GPU memory. Inference-mode zero-order optimization requires approximately 60 GB under the same conditions because it avoids stored activations, gradients, and optimizer states (Chaubard et al., 29 Sep 2026).

Let θ∈Rd\theta\in\mathbb{R}^d denote the parameter vector and L(θ)L(\theta) the loss. A generic zero-order update is

θk+1=θk−λg^(θk),\theta_{k+1}=\theta_k-\lambda \hat g(\theta_k),

where θ−ϵzi\theta-\epsilon z_i0 is inferred from loss evaluations rather than backpropagated derivatives.

Coordinate-wise finite differences require approximately θ−ϵzi\theta-\epsilon z_i1 function evaluations for a θ−ϵzi\theta-\epsilon z_i2-dimensional gradient. SPSA instead perturbs all parameters simultaneously. For Rademacher probes θ−ϵzi\theta-\epsilon z_i3, ordinary 1SPSA uses

θ−ϵzi\theta-\epsilon z_i4

Since θ−ϵzi\theta-\epsilon z_i5 coordinatewise, each probe requires two loss evaluations. The method’s evaluation cost is independent of the parameter dimension at the estimator level, although the total cost depends on the number of probes, accumulation steps, and forward evaluations.

The paper’s central computational thesis is to allocate a fixed forward-pass budget to larger effective batches and more perturbation probes, while using fewer sequential optimization steps. Probe evaluations can be distributed across devices, reducing sequential dependence even when the number of forward passes per step is large.

3. Clean evaluation and directional curvature

At an optimization iterate θ−ϵzi\theta-\epsilon z_i6, 1.5-SPSA first computes the clean loss

θ−ϵzi\theta-\epsilon z_i7

For each probe θ−ϵzi\theta-\epsilon z_i8, it evaluates

θ−ϵzi\theta-\epsilon z_i9

The clean evaluation is shared across all probes. It supplies the center point required for the second-order finite difference

L(θ)L(\theta)0

The associated first-order directional signal is

L(θ)L(\theta)1

The curvature estimate is a scalar associated with probe L(θ)L(\theta)2. It approximates L(θ)L(\theta)3, where L(θ)L(\theta)4, but it does not identify individual Hessian entries or a parameter-space diagonal.

Under three-times continuous differentiability, the finite-difference curvature has a third-order remainder. The paper writes the approximation in the form

L(θ)L(\theta)5

with L(θ)L(\theta)6 and L(θ)L(\theta)7. Consequently, if the third directional derivative is bounded by L(θ)L(\theta)8, then

L(θ)L(\theta)9

Thus, reducing c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.0 decreases finite-difference bias under smoothness assumptions, although excessively small c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.1 can make loss differences statistically or numerically noisy.

4. Probe-space preconditioning

A direct inverse-curvature weight c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.2 is unsafe because near-zero curvature can produce arbitrarily large weights, large curvature can excessively suppress updates, and neural-network curvature can be indefinite. 1.5-SPSA therefore uses the alpha-saturated weight

c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.3

The experiments use c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.4, while c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.5 controls the strength of curvature reweighting. At c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.6, the method approximately recovers standard 1SPSA; at c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.7, it applies full inverse-magnitude curvature scaling in probe space; intermediate values produce sublinear scaling. The default is c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.8.

The use of c^i=L(θ+ϵzi)−2L(θ)+L(θ−ϵzi)ϵ2.\hat c_i = \frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)} {\epsilon^2}.9 treats positive and negative curvature through magnitude rather than reversing the update direction. With LL0, a curvature magnitude of LL1 gives LL2, producing moderate rather than extreme suppression.

Aggregating LL3 probes, the update direction is

LL4

The parameter update is

LL5

with the paper tying LL6. The compact probe-space representation uses

LL7

where LL8 contains the probe vectors, LL9 is diagonal with entries c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.0, and c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.1 contains the central loss differences. The preconditioner therefore reweights sampled directions before their contributions are accumulated in parameter space.

This construction does not provide a full inverse-Hessian approximation. It is scalar per probe, depends on finite-difference curvature, is restricted to the span of the sampled probes, and uses a regularized fractional power rather than an exact inverse. The Johnson–Lindenstrauss discussion in the paper motivates the plausibility of preserving useful directional geometry in a lower-dimensional probe space, but it does not establish a Newton-like convergence theorem for nonconvex neural-network training.

5. Algorithmic procedure and computational cost

A 1.5-SPSA optimization step consists of the following operations:

  1. Select the probe radius c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.2, tied learning rate c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.3, accumulation count, probe count, saturation exponent c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.4, and regularization floor c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.5.
  2. Generate independent seeds defining Rademacher probes c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.6.
  3. Evaluate the clean loss c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.7.
  4. Evaluate c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.8 and c^i≈zi⊤∇2L(θ)zi.\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.9 for every probe.
  5. Compute each directional curvature $3/2$0.
  6. Compute each directional slope signal $3/2$1.
  7. Compute the curvature weight $3/2$2.
  8. Aggregate the weighted probe updates and apply the parameter update.

For 1SPSA, the paper defines the forward-pass cost as

$3/2$3

where $3/2$4 is the number of optimization steps and $3/2$5 is the number of accumulation steps. For 1.5-SPSA, the clean evaluation adds one batch evaluation per optimization step:

$3/2$6

The extra evaluation is shared by all probes and may already be required for tracking training or validation loss. The principal cost trade-off is therefore not necessarily fewer raw forward passes per step, but fewer sequential steps with more parallel probes.

The paper uses probe counts including $3/2$7, $3/2$8, and $3/2$9 in its main large-language-model configuration and explores θ∈Rd\theta\in\mathbb{R}^d0, θ∈Rd\theta\in\mathbb{R}^d1, θ∈Rd\theta\in\mathbb{R}^d2, and θ∈Rd\theta\in\mathbb{R}^d3 in an OPT-13B budget sweep. Perturbation-estimation variance is reported as approximately θ∈Rd\theta\in\mathbb{R}^d4, while batch-induced variance is approximately θ∈Rd\theta\in\mathbb{R}^d5 for effective batch size θ∈Rd\theta\in\mathbb{R}^d6.

The method uses gradient accumulation to obtain effective batch sizes of 128 or 256 in standard configurations, with the OPT-13B sweep reaching effective batch sizes up to 1024. Probe evaluations are parallelized across GPUs. An eight-GPU implementation distributes seeds and perturbation evaluations, gathers scalar losses, and synchronizes the model once per optimization step.

6. Hyperparameters, implementation, and empirical findings

The paper ties the probe radius and learning rate:

θ∈Rd\theta\in\mathbb{R}^d7

The standard sweep uses

θ∈Rd\theta\in\mathbb{R}^d8

and reports that θ∈Rd\theta\in\mathbb{R}^d9 is often effective. A plateau schedule halves both quantities after ten consecutive evaluations without validation-loss improvement. The default curvature parameters are L(θ)L(\theta)0 and L(θ)L(\theta)1. In the reported alpha ablation, L(θ)L(\theta)2 diverged, whereas L(θ)L(\theta)3, L(θ)L(\theta)4, L(θ)L(\theta)5, L(θ)L(\theta)6, and L(θ)L(\theta)7 were stable; the best listed test accuracy was 94.5% at L(θ)L(\theta)8.

The implementation represents Rademacher probes using one bit per parameter. For OPT-13B, the paper reports an unpacked bf16 perturbation of approximately 1.6 GB versus approximately 100 MB for a packed one-bit representation. Triton fused kernels unpack the bits, convert them to signs, scale them, and add them directly to model parameters or update buffers. The reported Triton bit-packed implementation achieved a 2.76-times end-to-end speedup over the original PyTorch implementation in the stated OPT-13B timing comparison.

On OPT-13B, 1.5-SPSA achieved 94.5% on SST-2, compared with 91.4% for MeZO and 92.0% for the reported BP baseline. The reported SST-2 comparison used approximately 179,000 forward passes and 70 steps for 1.5-SPSA, versus approximately 200,000 forward passes and 100,000 steps for MeZO. On RTE, BoolQ, WSC, and WiC, the reported 1.5-SPSA accuracies were 77.7%, 76.5%, 71.2%, and 61.9%, respectively.

For OPT-30B, the reported 1.5-SPSA accuracies were 94.5% on SST-2, 77.0% on RTE, 74.0% on BoolQ, 67.5% on WSC, and 59.3% on WiC. For Qwen3-8B, the corresponding values were 94.7%, 88.0%, 86.1%, 80.8%, and 71.2%. These results are empirical observations rather than general guarantees.

On a stiff paraboloid with condition numbers ranging from L(θ)L(\theta)9 to θk+1=θk−λg^(θk),\theta_{k+1}=\theta_k-\lambda \hat g(\theta_k),0, 1.5-SPSA approximately matched 1SPSA in well-conditioned cases and became increasingly faster as conditioning worsened, with up to approximately seven-times faster convergence. In a Differentiable Neural Computer stress test spanning approximately 300,000 to 1.1 billion parameters, 1.5-SPSA generally required fewer steps than 1SPSA, sometimes by approximately six times. Backpropagation was faster in many forward-pass-equivalent comparisons, but could not run the 1.1-billion-parameter model under the stated memory constraints.

The principal limitations are the high forward-pass cost, sensitivity to θk+1=θk−λg^(θk),\theta_{k+1}=\theta_k-\lambda \hat g(\theta_k),1, θk+1=θk−λg^(θk),\theta_{k+1}=\theta_k-\lambda \hat g(\theta_k),2, probe count, batch size, and θk+1=θk−λg^(θk),\theta_{k+1}=\theta_k-\lambda \hat g(\theta_k),3, finite-difference bias, and incomplete curvature normalization across batches or probes. The method does not maintain Adam-style momentum or long-term adaptive state. Its benefits are strongest in memory-limited and highly parallel settings, particularly when directional curvature varies substantially. The paper does not provide a complete nonconvex convergence theorem for deep neural-network training, and its backpropagation baselines were not exhaustively retuned for the large-batch, few-step regime used by 1.5-SPSA. Consequently, 1.5-SPSA is best characterized as 1SPSA plus a shared-center, directional-curvature correction that suppresses unstable high-curvature probes while preserving inference-mode memory usage.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to 1.5-SPSA.