---
title: 1.5-SPSA Optimization Method
url: https://www.emergentmind.com/topics/1-5-spsa
type: topic
---

# 1.5-SPSA Optimization Method

1.5-SPSA is a zero-order, inference-mode optimization method introduced for large-scale neural-network training. It extends ordinary one-probe SPSA (1SPSA) with one additional clean, unperturbed forward evaluation per optimization step. This shared center-point evaluation enables estimation of directional curvature for each random perturbation, which is then used to reweight probe contributions. The method does not construct or store a parameter-space Hessian; its preconditioner operates in the space of random perturbation directions [2609.38095].

## 1. Terminology and conceptual position

The name “1.5-SPSA” denotes an intermediate construction between first-order SPSA and a full second-order method. Ordinary two-sided SPSA estimates directional slopes using evaluations at $\theta+\epsilon z_i$ and $\theta-\epsilon z_i$. 1.5-SPSA adds the evaluation $L(\theta)$, producing the three-point stencil

$$
\hat c_i =
\frac{L(\theta+\epsilon z_i)-2L(\theta)+L(\theta-\epsilon z_i)}
{\epsilon^2}.
$$

For twice-differentiable $L$, this quantity approximates the curvature along the probe direction:

$$
\hat c_i \approx z_i^\top \nabla^2L(\theta)z_i.
$$

The method is therefore “one-and-a-half” in the sense that it retains SPSA’s first-order random-probe structure while adding a limited second-order correction. It is not a Newton method, does not estimate a full Hessian, and does not compute a diagonal Hessian in parameter coordinates.

The method should be distinguished from several unrelated uses of “1.5-SPSA.” The SPS upgrade paper does not define the term, and the QAOA, beamforming, SPSA-theory, and VQE papers likewise do not introduce a method with this name. In particular, the $3/2$ exponent appearing in finite-shot measurement-complexity analyses is not a fractional-order SPSA algorithm. Guided-SPSA combines parameter-shift gradients and SPSA but is a sample-partitioned hybrid rather than 1.5-SPSA. The method defined here is specifically probe-space preconditioning for zero-order training [1409.5821; 2104.09199; 2404.15751; 2608.09810].

## 2. Motivation and optimization setting

Backpropagation requires storing intermediate activations and, when used with Adam, maintaining first- and second-moment optimizer states. The paper gives training OPT-30B with Adam, batch size 8, and sequence length 2048 as requiring approximately 600 GB of GPU memory. Inference-mode zero-order optimization requires approximately 60 GB under the same conditions because it avoids stored activations, gradients, and optimizer states [2609.38095].

Let $\theta\in\mathbb{R}^d$ denote the parameter vector and $L(\theta)$ the loss. A generic zero-order update is

$$
\theta_{k+1}=\theta_k-\lambda \hat g(\theta_k),
$$

where $\hat g$ is inferred from loss evaluations rather than backpropagated derivatives.

Coordinate-wise finite differences require approximately $2d$ function evaluations for a $d$-dimensional gradient. SPSA instead perturbs all parameters simultaneously. For Rademacher probes $z_i\in\{-1,+1\}^d$, ordinary 1SPSA uses

$$
\hat g(\theta)
=
\frac{1}{n_{\mathrm{pert}}}
\sum_{i=1}^{n_{\mathrm{pert}}}
\frac{L(\theta+\epsilon z_i)-L(\theta-\epsilon z_i)}
{2\epsilon}\,z_i^{-1}.
$$

Since $z_i^{-1}=z_i$ coordinatewise, each probe requires two loss evaluations. The method’s evaluation cost is independent of the parameter dimension at the estimator level, although the total cost depends on the number of probes, accumulation steps, and forward evaluations.

The paper’s central computational thesis is to allocate a fixed forward-pass budget to larger effective batches and more perturbation probes, while using fewer sequential optimization steps. Probe evaluations can be distributed across devices, reducing sequential dependence even when the number of forward passes per step is large.

## 3. Clean evaluation and directional curvature

At an optimization iterate $\theta$, 1.5-SPSA first computes the clean loss

$$
L_0=L(\theta).
$$

For each probe $z_i$, it evaluates

$$
L_{+,i}=L(\theta+\epsilon z_i),
\qquad
L_{-,i}=L(\theta-\epsilon z_i).
$$

The clean evaluation is shared across all probes. It supplies the center point required for the second-order finite difference

$$
\hat c_i =
\frac{L_{+,i}-2L_0+L_{-,i}}{\epsilon^2}.
$$

The associated first-order directional signal is

$$
g_i=L_{+,i}-L_{-,i}.
$$

The curvature estimate is a scalar associated with probe $z_i$. It approximates $z_i^\top H z_i$, where $H=\nabla^2L(\theta)$, but it does not identify individual Hessian entries or a parameter-space diagonal.

Under three-times continuous differentiability, the finite-difference curvature has a third-order remainder. The paper writes the approximation in the form

$$
\hat c(z)
=
z^\top Hz+
\frac{\epsilon}{6}
\left(
\nabla^3L(\theta+\xi_+\epsilon z)[z,z,z]
-
\nabla^3L(\theta+\xi_-\epsilon z)[z,z,z]
\right),
$$

with $\xi_+\in(0,1)$ and $\xi_-\in(-1,0)$. Consequently, if the third directional derivative is bounded by $M\|z\|^3$, then

$$
|\hat c(z)-z^\top Hz|
\leq
\frac{\epsilon}{3}M\|z\|^3.
$$

Thus, reducing $\epsilon$ decreases finite-difference bias under smoothness assumptions, although excessively small $\epsilon$ can make loss differences statistically or numerically noisy.

## 4. Probe-space preconditioning

A direct inverse-curvature weight $1/\hat c_i$ is unsafe because near-zero curvature can produce arbitrarily large weights, large curvature can excessively suppress updates, and neural-network curvature can be indefinite. 1.5-SPSA therefore uses the alpha-saturated weight

$$
w_i=
\frac{1}
{\max\left(\lambda_{\mathrm{reg}},|\hat c_i|^\alpha\right)}.
$$

The experiments use $\lambda_{\mathrm{reg}}=1$, while $\alpha\in[0,1]$ controls the strength of curvature reweighting. At $\alpha=0$, the method approximately recovers standard 1SPSA; at $\alpha=1$, it applies full inverse-magnitude curvature scaling in probe space; intermediate values produce sublinear scaling. The default is $\alpha=0.1$.

The use of $|\hat c_i|$ treats positive and negative curvature through magnitude rather than reversing the update direction. With $\alpha=0.1$, a curvature magnitude of $10^8$ gives $|\hat c_i|^{0.1}\approx6.3$, producing moderate rather than extreme suppression.

Aggregating $n=n_{\mathrm{pert}}$ probes, the update direction is

$$
\Delta\theta
=
\frac{1}{2n}
\sum_{i=1}^{n}
\left(
\frac{L(\theta+\epsilon z_i)-L(\theta-\epsilon z_i)}
{\max(\lambda_{\mathrm{reg}},|\hat c_i|^\alpha)}
\right)z_i.
$$

The parameter update is

$$
\theta\leftarrow\theta-\lambda\Delta\theta,
$$

with the paper tying $\lambda=\epsilon$. The compact probe-space representation uses

$$
\Delta\theta=-\eta ZW\Delta\ell,
$$

where $Z$ contains the probe vectors, $W$ is diagonal with entries $w_i$, and $\Delta\ell$ contains the central loss differences. The preconditioner therefore reweights sampled directions before their contributions are accumulated in parameter space.

This construction does not provide a full inverse-Hessian approximation. It is scalar per probe, depends on finite-difference curvature, is restricted to the span of the sampled probes, and uses a regularized fractional power rather than an exact inverse. The Johnson–Lindenstrauss discussion in the paper motivates the plausibility of preserving useful directional geometry in a lower-dimensional probe space, but it does not establish a Newton-like convergence theorem for nonconvex neural-network training.

## 5. Algorithmic procedure and computational cost

A 1.5-SPSA optimization step consists of the following operations:

1. Select the probe radius $\epsilon$, tied learning rate $\lambda=\epsilon$, accumulation count, probe count, saturation exponent $\alpha$, and regularization floor $\lambda_{\mathrm{reg}}$.
2. Generate independent seeds defining Rademacher probes $z_i\in\{-1,+1\}^d$.
3. Evaluate the clean loss $L_0=L(\theta)$.
4. Evaluate $L_{+,i}$ and $L_{-,i}$ for every probe.
5. Compute each directional curvature $\hat c_i$.
6. Compute each directional slope signal $g_i$.
7. Compute the curvature weight $w_i$.
8. Aggregate the weighted probe updates and apply the parameter update.

For 1SPSA, the paper defines the forward-pass cost as

$$
F_{\mathrm{1SPSA}}
=
s\,a\,2n_{\mathrm{pert}},
$$

where $s$ is the number of optimization steps and $a$ is the number of accumulation steps. For 1.5-SPSA, the clean evaluation adds one batch evaluation per optimization step:

$$
F_{\mathrm{1.5SPSA}}
=
s\,a\,(2n_{\mathrm{pert}}+1).
$$

The extra evaluation is shared by all probes and may already be required for tracking training or validation loss. The principal cost trade-off is therefore not necessarily fewer raw forward passes per step, but fewer sequential steps with more parallel probes.

The paper uses probe counts including $40$, $60$, and $100$ in its main large-language-model configuration and explores $20$, $40$, $160$, and $640$ in an OPT-13B budget sweep. Perturbation-estimation variance is reported as approximately $O(1/n_{\mathrm{pert}})$, while batch-induced variance is approximately $O(1/B)$ for effective batch size $B$.

The method uses gradient accumulation to obtain effective batch sizes of 128 or 256 in standard configurations, with the OPT-13B sweep reaching effective batch sizes up to 1024. Probe evaluations are parallelized across GPUs. An eight-GPU implementation distributes seeds and perturbation evaluations, gathers scalar losses, and synchronizes the model once per optimization step.

## 6. Hyperparameters, implementation, and empirical findings

The paper ties the probe radius and learning rate:

$$
\lambda=\epsilon.
$$

The standard sweep uses

$$
\lambda=\epsilon
\in
\{10^{-3},5\cdot10^{-4},\ldots,10^{-7}\},
$$

and reports that $\lambda=\epsilon=10^{-4}$ is often effective. A plateau schedule halves both quantities after ten consecutive evaluations without validation-loss improvement. The default curvature parameters are $\lambda_{\mathrm{reg}}=1$ and $\alpha=0.1$. In the reported alpha ablation, $\alpha=0.01$ diverged, whereas $\alpha=0.1$, $0.25$, $0.5$, $0.75$, and $1.0$ were stable; the best listed test accuracy was 94.5% at $\alpha=0.1$.

The implementation represents Rademacher probes using one bit per parameter. For OPT-13B, the paper reports an unpacked bf16 perturbation of approximately 1.6 GB versus approximately 100 MB for a packed one-bit representation. Triton fused kernels unpack the bits, convert them to signs, scale them, and add them directly to model parameters or update buffers. The reported Triton bit-packed implementation achieved a 2.76-times end-to-end speedup over the original PyTorch implementation in the stated OPT-13B timing comparison.

On OPT-13B, 1.5-SPSA achieved 94.5% on SST-2, compared with 91.4% for MeZO and 92.0% for the reported BP baseline. The reported SST-2 comparison used approximately 179,000 forward passes and 70 steps for 1.5-SPSA, versus approximately 200,000 forward passes and 100,000 steps for MeZO. On RTE, BoolQ, WSC, and WiC, the reported 1.5-SPSA accuracies were 77.7%, 76.5%, 71.2%, and 61.9%, respectively.

For OPT-30B, the reported 1.5-SPSA accuracies were 94.5% on SST-2, 77.0% on RTE, 74.0% on BoolQ, 67.5% on WSC, and 59.3% on WiC. For Qwen3-8B, the corresponding values were 94.7%, 88.0%, 86.1%, 80.8%, and 71.2%. These results are empirical observations rather than general guarantees.

On a stiff paraboloid with condition numbers ranging from $1$ to $1000$, 1.5-SPSA approximately matched 1SPSA in well-conditioned cases and became increasingly faster as conditioning worsened, with up to approximately seven-times faster convergence. In a Differentiable Neural Computer stress test spanning approximately 300,000 to 1.1 billion parameters, 1.5-SPSA generally required fewer steps than 1SPSA, sometimes by approximately six times. Backpropagation was faster in many forward-pass-equivalent comparisons, but could not run the 1.1-billion-parameter model under the stated memory constraints.

The principal limitations are the high forward-pass cost, sensitivity to $\epsilon$, $\lambda$, probe count, batch size, and $\alpha$, finite-difference bias, and incomplete curvature normalization across batches or probes. The method does not maintain Adam-style momentum or long-term adaptive state. Its benefits are strongest in memory-limited and highly parallel settings, particularly when directional curvature varies substantially. The paper does not provide a complete nonconvex convergence theorem for deep neural-network training, and its backpropagation baselines were not exhaustively retuned for the large-batch, few-step regime used by 1.5-SPSA. Consequently, 1.5-SPSA is best characterized as 1SPSA plus a shared-center, directional-curvature correction that suppresses unstable high-curvature probes while preserving inference-mode memory usage.

Source: https://www.emergentmind.com/topics/1-5-spsa