Probe-Space Preconditioning for Fast and Stable Zero-Order Training
This presentation explores a breakthrough in training large neural networks without backpropagation. The paper introduces 1.5-SPSA, a zero-order optimization method that combines aggressive forward-pass aggregation with probe-space curvature correction to achieve competitive accuracy in just tens of optimization steps while using only 10% of the memory required by traditional backpropagation. The method demonstrates particularly strong results on ill-conditioned post-training objectives for billion-parameter language models.Script
Training a 30 billion parameter language model with backpropagation requires roughly 600 gigabytes of GPU memory. Zero-order optimization cuts that to 60 gigabytes by eliminating stored gradients entirely, but the trade-off has always been thousands of wasted function evaluations and unstable convergence.
The authors propose 1.5-SPSA, which adds a single clean loss evaluation to estimate directional curvature along each random probe. Instead of treating all search directions equally, the method downweights directions with extreme curvature, stabilizing updates in the chaotic loss landscapes of post-training tasks. This operates in probe space, not parameter space, so it requires no Hessian storage and no per-parameter optimizer state.
The first key insight is about compute allocation. Rather than spreading a fixed budget across thousands of cheap optimization steps, the method invests forward passes into larger effective batches and more perturbation directions per step. On OPT-13B fine-tuning, this permits a learning rate 500 times larger than prior zero-order methods and convergence in just 80 steps instead of 100,000.
The authors measured directional curvature on Qwen 8B and found variation across eight orders of magnitude, with both strongly positive and strongly negative regions. Applying a single step size to such landscapes guarantees instability. On a controlled stiff paraboloid, 1.5-SPSA converged in 6 steps where standard SPSA required 2,000, with speedups growing proportionally to the condition number.
On OPT-13B across five language tasks, 1.5-SPSA reached 94.5% on sentiment classification and 77.7% on textual entailment, exceeding both prior zero-order baselines and the reported backpropagation result. However, the gains were not uniform. On some tasks the method matched but did not surpass backpropagation, revealing that curvature correction improves optimization reliability but does not universally improve generalization.
The method opens a broader corridor of stable hyperparameters and proves that inference-mode optimization can compete with backpropagation in highly ill-conditioned settings. If you want to explore the technical details or generate your own video summary, visit EmergentMind.com and see what probe-space thinking can unlock.