---
title: Randomized Directional Smoothing Techniques
url: https://www.emergentmind.com/topics/randomized-directional-smoothing
type: topic
---

# Randomized Directional Smoothing Techniques

Randomized directional smoothing encompasses a family of techniques in optimization, control, robustness certification, and model security, exploiting directional randomization and averaging to smooth non-smooth functions, stabilize optimization, or increase model robustness. These methods generalize classical randomized smoothing by replacing isotropic perturbations with structured, often directionally-aligned, noise. The canonical instantiations include randomized zeroth-order optimization with orthogonal directions, directional embedding smoothing for model defense, anisotropic Gaussian smoothing for accelerated convergence, and parameter-space smoothing for certifiable robustness. These techniques share the principle of Monte Carlo approximation of non-local quantities using random perturbations along selected directions or distributional axes.

## 1. Mathematical Foundations of Randomized Directional Smoothing

Randomized directional smoothing replaces a potentially non-smooth or brittle function $f:\mathbb{R}^d\to\mathbb{R}$ by a locally averaged (smoothed) version along chosen directions. The canonical scalar-valued smoothing takes the form
$$
f_\Sigma(x) = \mathbb{E}_{u \sim \mathcal{N}(0, \Sigma)}[f(x + u)]
$$
where $\Sigma$ is a symmetric, positive definite covariance matrix that determines the smoothing geometry. Isotropic smoothing takes $\Sigma = \sigma^2 I$; directional or anisotropic smoothing selects $\Sigma$ with nonuniform eigenvalues or aligns noise with specific directions.

In the context of zeroth-order gradient estimation, one typically draws random directions $\{u_i\}_{i=1}^m$ forming an orthonormal basis and constructs a Monte Carlo gradient approximation:
$$
g_h(x) = \frac{1}{h}\sum_{i=1}^{m} [f(x + h u_i) - f(x)] u_i
$$
with $h > 0$ the smoothing radius. For the general anisotropic setting, the gradient of the smoothed function can be written as
$$
\nabla f_\Sigma(x) = \Sigma^{-1} \mathbb{E}_{u \sim \mathcal{N}(0,\Sigma)}[u f(x+u)]
$$
as established in anisotropic smoothing frameworks [2411.11747].

Directional embedding smoothing for vision-language models perturbs token embeddings $e \in \mathbb{R}^d$ by adding noise of the form
$$
e' = \left(1 + \frac{z}{\|e\|_2}\right) e, \quad z \sim \mathcal{N}(0, \sigma^2 d)
$$
which ensures noise is injected strictly along the embedding direction [2603.15259].

## 2. Algorithmic Instantiations and Implementations

Several algorithmic schemes operationalize randomized directional smoothing, depending on the domain and goal:

- **Zeroth-order Optimization with Orthogonal Directions:** At each iteration, draw $m$ random orthonormal vectors $\{u_i\}$ and compute finite-difference estimates in these directions. The estimate is lifted to the full space by a linear map. This underpins the Randomized Directional Smoothing (RDS) and spherical smoothing for optimization [2107.03941].

- **Anisotropic Gaussian Smoothing (AGS):** Build the smoothing covariance $\Sigma$ adaptive to curvature or recent gradients. Monte Carlo estimators of $\nabla f_\Sigma(x)$ are used in AGS-GD, AGS-SGD, and AGS-Adam algorithms. Adaptation strategies leverage local Hessian or gradient moments to orient smoothing along high-curvature directions [2411.11747].

- **Directional Embedding Smoothing in RESTA:** For robust vision-language models, perturb each token embedding directionally as above, generate $k$ noisy embedding sequences, autoregressively decode in parallel, and select next tokens by majority vote. The smoothing relies on injection of Gaussian noise strictly aligned with each token vector [2603.15259].

- **Randomized Smoothing for Control and Certification:** In optimal control of nonsmooth systems, perturbed states are averaged to define smooth surrogate dynamics amenable to gradient-based methods. In image certification, transformation parameters are smoothed with Gaussian noise, and error bounds are established in parameter space for robust prediction [2203.03986, 2002.12463].

The following table summarizes core algorithmic settings:

| Application Domain                | Smoothing Directionality     | Reference        |
|-----------------------------------|-----------------------------|------------------|
| Zeroth-order Optimization         | Random orthonormal basis    | [2107.03941]     |
| Anisotropic Gradient-based Opt.   | Adaptive covariance (Σ)     | [2411.11747]     |
| Vision-Language Model Defense     | Embedding vector direction  | [2603.15259]     |
| Control/Certification             | State or parameter space    | [2203.03986], [2002.12463] |

## 3. Theoretical Properties and Convergence Analyses

Randomized directional smoothing introduces bias and variance trade-offs that are quantifiable both for estimation accuracy and optimization convergence:

- For RDS in optimization, the estimator bias is $O(h)$ and variance is $O(\frac{d}{m} \|\nabla f(x)\|^2 + \frac{L^2 d^2}{4m} h^2)$. Convergence rates for convex $L$-smooth objectives under appropriate step-size and smoothing radius yield $\mathbb{E}[f(\hat x_K) - f_*] \leq O(1/K) + O(h)$, with linear convergence to a neighborhood under the Polyak-Łojasiewicz (PL) condition [2107.03941].

- In AGS methods, rates in both convex and nonconvex settings obtain, with convergence balls' radii dictated by $\|\Sigma\|, \|\Sigma^{-1}\|$, and step-sizes. Anisotropic smoothing recovers classical isotropic bounds in the special case $\Sigma = \sigma^2 I$ but often leads to improved empirical convergence, particularly in ill-conditioned or ridge-like landscapes [2411.11747].

- For RESTA with directional embedding smoothing, no formal robustness certificates are proved for vision-language models; effectiveness is motivated by the disruption of local adversarial paths in embedding space [2603.15259].

- In randomized smoothing-based certification for parameterized transformations, explicit certification radii in parameter space are derived, guaranteeing robustness to geometric perturbations up to a provable threshold [2002.12463].

## 4. Practical Considerations and Hyperparameter Selection

Effective application of randomized directional smoothing relies on careful tuning of parameters:

- **Number of Random Directions / Samples ($m, M, k$):** $m$ balances estimator variance and per-iteration cost. In high dimensions, $m \ll d$ is common for optimization, while for model defense $k=10$ suffices for majority voting [2107.03941, 2603.15259].

- **Noise Scale ($h, \sigma, \|\Sigma\|$):** The smoothing radius determines bias-variance trade-off. Too small $h$ increases variance; too large introduces excessive bias. Practical heuristics select $h \approx 10^{-3} \|\nabla f\|$ for optimization, or $0.2 \leq \sigma \leq 0.6$ for embeddings in Gemma [2411.11747, 2603.15259].

- **Covariance Design (Anisotropy):** For AGS, $\Sigma$ is often built using local Hessian or moving gradient moments to adapt directionality [2411.11747].

- **Algorithmic Stability:** Aggressive adaptation of $\Sigma$ can violate smoothness assumptions. Conservative updates and periodic adaptation are recommended [2411.11747].

- **System Deployment:** In robust inference (e.g., RESTA), randomized directional smoothing acts as a lightweight, inference-time layer. Composability with other defenses is possible but does not obviate the need for broader system-level security, including red teaming and alignment training [2603.15259].

## 5. Empirical Results and Applications

Empirical evidence demonstrates the utility of randomized directional smoothing across domains:

- **Model Robustness and Security:** On the JailBreakV-28K suite, RESTA with directional noise reduced attack success rate on LLaVA-1.5-7B from 50.13% to 25.93% with only a minor utility loss (ScienceQA score drop from 64.07% to 61.42%). Isotropic noise failed to obtain comparable gains [2603.15259].

- **Optimization:** RDS achieves convergence rates $O(1/K)$ for convex and linear up to $O(h^2)$ for PL objectives. Empirically, moderate subspace dimension $m \approx d/10$ provides best trade-off between cost and convergence [2107.03941].

- **Control:** In optimal control of systems with non-smooth dynamics (Coulomb friction, impacts), randomized smoothing enables second-order methods (R-DDP) to succeed where classical methods or pure RL either fail or require excessive samples. R-DDP matched state-of-the-art RL with $10^1$–$10^2$ times fewer samples on contact-rich tasks [2203.03986].

- **Certification:** Robustness certificates against geometric transformations were established for smoothing in parameter space, providing high-confidence intervals and individual certifiable radii for image classification [2002.12463].

## 6. Variants, Special Cases, and Extensions

Randomized directional smoothing admits several reductions and related methodologies:

- **Spherical Smoothing:** $m=1$ recovers classical two-point estimation along a single random direction [2107.03941].
- **Coordinate Descent:** Employing coordinate axes as directions recovers randomized coordinate descent.
- **Anisotropic vs. Isotropic Smoothing:** Taking $\Sigma = \sigma^2 I$ yields isotropic smoothing (classical randomized smoothing as in Cohen et al.), while general $\Sigma$ enables highly directional or adaptive smoothing [2411.11747].
- **Embedding Smoothing vs. Parameter Space Smoothing:** In vision-language models, smoothing is performed in the embedding space; in geometric robustness it is performed in transformation parameter space [2603.15259, 2002.12463].

A plausible implication is that further adaptation of the covariance or smoothing geometry—guided by application-specific priors or learned structural information—may further enhance the security-utility or convergence trade-off.

## 7. Limitations and Future Directions

While randomized directional smoothing significantly advances optimization efficiency, robustness, and model security, several limitations and open questions persist:

- **Certification Gaps:** While formal certificates exist in parameter-space smoothing, defenses like RESTA-directional remain heuristic, lacking formal robustness or failure bounds in adversarial settings [2603.15259].

- **Adversarial Adaptivity:** Defensive smoothing may only temporarily mitigate attack success until adaptive adversaries circumvent the smoothed geometry.

- **Compositional Effects:** How randomized smoothing interacts with system-level composition—multiple defense layers, alignment training, and red teaming—remains an area for empirical and theoretical investigation [2603.15259].

Future work is expected to close the certification gaps for adversarial robustness, develop more sophisticated mechanisms for adaptive smoothing geometry, and establish compositional guarantees in multi-layer robust learning systems.

Source: https://www.emergentmind.com/topics/randomized-directional-smoothing