1.5-SPSA Optimization Method
- 1.5-SPSA is a zero-order optimization method specifically designed for neural-network training, characterized by adding a central clean evaluation per optimization step to estimate and use directional curvature.
- 1.5-SPSA achieves improved performance by reweighting gradient directions based on curvature estimates, which enhances the efficiency of neural network training, particularly in memory-constrained environments. It requires fewer forward passes per step but may be sensitive to hyperparameter settings.
- This method parallelizes probe evaluations across GPUs to optimize training time, effectively treating positive and negative curvature magnitudes equally, and resulting in moderate suppression of extreme curvatures, which retains inference-mode retention characteristics in training.
1.5-SPSA is a zero-order, inference-mode optimization method introduced for large-scale neural-network training. It extends ordinary one-probe SPSA (1SPSA) with one additional clean, unperturbed forward evaluation per optimization step. This shared center-point evaluation enables estimation of directional curvature for each random perturbation, which is then used to reweight probe contributions. The method does not construct or store a parameter-space Hessian; its preconditioner operates in the space of random perturbation directions (Chaubard et al., 29 Sep 2026).
1. Terminology and conceptual position
The name “1.5-SPSA” denotes an intermediate construction between first-order SPSA and a full second-order method. Ordinary two-sided SPSA estimates directional slopes using evaluations at and . 1.5-SPSA adds the evaluation , producing the three-point stencil
For twice-differentiable , this quantity approximates the curvature along the probe direction:
The method is therefore “one-and-a-half” in the sense that it retains SPSA’s first-order random-probe structure while adding a limited second-order correction. It is not a Newton method, does not estimate a full Hessian, and does not compute a diagonal Hessian in parameter coordinates.
The method should be distinguished from several unrelated uses of “1.5-SPSA.” The SPS upgrade paper does not define the term, and the QAOA, beamforming, SPSA-theory, and VQE papers likewise do not introduce a method with this name. In particular, the $3/2$ exponent appearing in finite-shot measurement-complexity analyses is not a fractional-order SPSA algorithm. Guided-SPSA combines parameter-shift gradients and SPSA but is a sample-partitioned hybrid rather than 1.5-SPSA. The method defined here is specifically probe-space preconditioning for zero-order training (Goddard et al., 2014, Chen et al., 2021, Periyasamy et al., 2024, Qin, 10 Aug 2026).
2. Motivation and optimization setting
Backpropagation requires storing intermediate activations and, when used with Adam, maintaining first- and second-moment optimizer states. The paper gives training OPT-30B with Adam, batch size 8, and sequence length 2048 as requiring approximately 600 GB of GPU memory. Inference-mode zero-order optimization requires approximately 60 GB under the same conditions because it avoids stored activations, gradients, and optimizer states (Chaubard et al., 29 Sep 2026).
Let denote the parameter vector and the loss. A generic zero-order update is
where 0 is inferred from loss evaluations rather than backpropagated derivatives.
Coordinate-wise finite differences require approximately 1 function evaluations for a 2-dimensional gradient. SPSA instead perturbs all parameters simultaneously. For Rademacher probes 3, ordinary 1SPSA uses
4
Since 5 coordinatewise, each probe requires two loss evaluations. The method’s evaluation cost is independent of the parameter dimension at the estimator level, although the total cost depends on the number of probes, accumulation steps, and forward evaluations.
The paper’s central computational thesis is to allocate a fixed forward-pass budget to larger effective batches and more perturbation probes, while using fewer sequential optimization steps. Probe evaluations can be distributed across devices, reducing sequential dependence even when the number of forward passes per step is large.
3. Clean evaluation and directional curvature
At an optimization iterate 6, 1.5-SPSA first computes the clean loss
7
For each probe 8, it evaluates
9
The clean evaluation is shared across all probes. It supplies the center point required for the second-order finite difference
0
The associated first-order directional signal is
1
The curvature estimate is a scalar associated with probe 2. It approximates 3, where 4, but it does not identify individual Hessian entries or a parameter-space diagonal.
Under three-times continuous differentiability, the finite-difference curvature has a third-order remainder. The paper writes the approximation in the form
5
with 6 and 7. Consequently, if the third directional derivative is bounded by 8, then
9
Thus, reducing 0 decreases finite-difference bias under smoothness assumptions, although excessively small 1 can make loss differences statistically or numerically noisy.
4. Probe-space preconditioning
A direct inverse-curvature weight 2 is unsafe because near-zero curvature can produce arbitrarily large weights, large curvature can excessively suppress updates, and neural-network curvature can be indefinite. 1.5-SPSA therefore uses the alpha-saturated weight
3
The experiments use 4, while 5 controls the strength of curvature reweighting. At 6, the method approximately recovers standard 1SPSA; at 7, it applies full inverse-magnitude curvature scaling in probe space; intermediate values produce sublinear scaling. The default is 8.
The use of 9 treats positive and negative curvature through magnitude rather than reversing the update direction. With 0, a curvature magnitude of 1 gives 2, producing moderate rather than extreme suppression.
Aggregating 3 probes, the update direction is
4
The parameter update is
5
with the paper tying 6. The compact probe-space representation uses
7
where 8 contains the probe vectors, 9 is diagonal with entries 0, and 1 contains the central loss differences. The preconditioner therefore reweights sampled directions before their contributions are accumulated in parameter space.
This construction does not provide a full inverse-Hessian approximation. It is scalar per probe, depends on finite-difference curvature, is restricted to the span of the sampled probes, and uses a regularized fractional power rather than an exact inverse. The Johnson–Lindenstrauss discussion in the paper motivates the plausibility of preserving useful directional geometry in a lower-dimensional probe space, but it does not establish a Newton-like convergence theorem for nonconvex neural-network training.
5. Algorithmic procedure and computational cost
A 1.5-SPSA optimization step consists of the following operations:
- Select the probe radius 2, tied learning rate 3, accumulation count, probe count, saturation exponent 4, and regularization floor 5.
- Generate independent seeds defining Rademacher probes 6.
- Evaluate the clean loss 7.
- Evaluate 8 and 9 for every probe.
- Compute each directional curvature $3/2$0.
- Compute each directional slope signal $3/2$1.
- Compute the curvature weight $3/2$2.
- Aggregate the weighted probe updates and apply the parameter update.
For 1SPSA, the paper defines the forward-pass cost as
$3/2$3
where $3/2$4 is the number of optimization steps and $3/2$5 is the number of accumulation steps. For 1.5-SPSA, the clean evaluation adds one batch evaluation per optimization step:
$3/2$6
The extra evaluation is shared by all probes and may already be required for tracking training or validation loss. The principal cost trade-off is therefore not necessarily fewer raw forward passes per step, but fewer sequential steps with more parallel probes.
The paper uses probe counts including $3/2$7, $3/2$8, and $3/2$9 in its main large-language-model configuration and explores 0, 1, 2, and 3 in an OPT-13B budget sweep. Perturbation-estimation variance is reported as approximately 4, while batch-induced variance is approximately 5 for effective batch size 6.
The method uses gradient accumulation to obtain effective batch sizes of 128 or 256 in standard configurations, with the OPT-13B sweep reaching effective batch sizes up to 1024. Probe evaluations are parallelized across GPUs. An eight-GPU implementation distributes seeds and perturbation evaluations, gathers scalar losses, and synchronizes the model once per optimization step.
6. Hyperparameters, implementation, and empirical findings
The paper ties the probe radius and learning rate:
7
The standard sweep uses
8
and reports that 9 is often effective. A plateau schedule halves both quantities after ten consecutive evaluations without validation-loss improvement. The default curvature parameters are 0 and 1. In the reported alpha ablation, 2 diverged, whereas 3, 4, 5, 6, and 7 were stable; the best listed test accuracy was 94.5% at 8.
The implementation represents Rademacher probes using one bit per parameter. For OPT-13B, the paper reports an unpacked bf16 perturbation of approximately 1.6 GB versus approximately 100 MB for a packed one-bit representation. Triton fused kernels unpack the bits, convert them to signs, scale them, and add them directly to model parameters or update buffers. The reported Triton bit-packed implementation achieved a 2.76-times end-to-end speedup over the original PyTorch implementation in the stated OPT-13B timing comparison.
On OPT-13B, 1.5-SPSA achieved 94.5% on SST-2, compared with 91.4% for MeZO and 92.0% for the reported BP baseline. The reported SST-2 comparison used approximately 179,000 forward passes and 70 steps for 1.5-SPSA, versus approximately 200,000 forward passes and 100,000 steps for MeZO. On RTE, BoolQ, WSC, and WiC, the reported 1.5-SPSA accuracies were 77.7%, 76.5%, 71.2%, and 61.9%, respectively.
For OPT-30B, the reported 1.5-SPSA accuracies were 94.5% on SST-2, 77.0% on RTE, 74.0% on BoolQ, 67.5% on WSC, and 59.3% on WiC. For Qwen3-8B, the corresponding values were 94.7%, 88.0%, 86.1%, 80.8%, and 71.2%. These results are empirical observations rather than general guarantees.
On a stiff paraboloid with condition numbers ranging from 9 to 0, 1.5-SPSA approximately matched 1SPSA in well-conditioned cases and became increasingly faster as conditioning worsened, with up to approximately seven-times faster convergence. In a Differentiable Neural Computer stress test spanning approximately 300,000 to 1.1 billion parameters, 1.5-SPSA generally required fewer steps than 1SPSA, sometimes by approximately six times. Backpropagation was faster in many forward-pass-equivalent comparisons, but could not run the 1.1-billion-parameter model under the stated memory constraints.
The principal limitations are the high forward-pass cost, sensitivity to 1, 2, probe count, batch size, and 3, finite-difference bias, and incomplete curvature normalization across batches or probes. The method does not maintain Adam-style momentum or long-term adaptive state. Its benefits are strongest in memory-limited and highly parallel settings, particularly when directional curvature varies substantially. The paper does not provide a complete nonconvex convergence theorem for deep neural-network training, and its backpropagation baselines were not exhaustively retuned for the large-batch, few-step regime used by 1.5-SPSA. Consequently, 1.5-SPSA is best characterized as 1SPSA plus a shared-center, directional-curvature correction that suppresses unstable high-curvature probes while preserving inference-mode memory usage.