Papers
Topics
Authors
Recent
Search
2000 character limit reached

Activation Perturbation for Exploration (APEX)

Updated 10 February 2026
  • The paper introduces APEX as a probing technique that injects Gaussian noise into hidden activations to interpolate between input-sensitive and model-driven behaviors.
  • APEX employs Monte Carlo sampling over multiple noise scales to compute escape noise, providing precise diagnostics for sample regularity, semantic alignment, and backdoor detection.
  • By modulating noise levels, APEX enables researchers to gain both local insights and global bias analysis without needing model retraining or ensemble methods.

Activation Perturbation for EXploration (APEX) is an inference-time probing paradigm for neural networks that systematically injects Gaussian noise into hidden activations, while holding both the model input and parameters fixed. APEX is designed to address limitations inherent in input-space and parameter perturbation approaches, providing a direct lens into the structure and regularities encoded in intermediate network representations. By varying the noise scale, APEX enables a controlled transition from sample-dependent, input-sensitive responses to model-driven, input-agnostic behaviors, offering both local and global perspectives on network decision processes (Ren et al., 3 Feb 2026).

1. Formalism and Algorithmic Specification

Consider an LL-layer feed-forward network fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c with pre-activations and post-activations given by

zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)

where θ=(W,b)\theta = (W, b). APEX introduces additive Gaussian noise to each post-activation at inference:

a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)

for each ℓ=1,…,L\ell = 1,\dots,L and noise scale σ>0\sigma > 0. The final logits are

s(x;σ)=U a~L(x;σ)+cs(x; \sigma) = U\,\tilde a_L(x; \sigma) + c

and the predicted label is

k∗(x;σ)=arg⁡max⁡ksk(x;σ)k^*(x;\sigma) = \arg\max_k s_k(x;\sigma)

The empirical output distribution is estimated via TT Monte Carlo forward passes:

fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c0

APEX thus interpolates between sample-dependent (fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c1) and model-dependent (fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c2) response regimes.

Algorithmically, for each input fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c3, each chosen noise scale fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c4, and each of fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c5 forward passes:

  • Independently sample fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c6 for all fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c7
  • Inject fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c8 after each layer’s activation in the network
  • Record the resulting top-1 class fθ:Rd→Rcf_\theta: \mathbb{R}^d \rightarrow \mathbb{R}^c9
  • Aggregate to yield zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)0

The escape noise zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)1 for input zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)2 is defined as the minimal zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)3 where the original prediction’s probability drops below a fixed threshold zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)4.

2. Theoretical Framework and Decomposition

At the core of APEX is a decomposition theorem applicable for any zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)5 with zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)6 and all zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)7:

zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)8

where zℓ=Wℓaℓ−1+bℓ,aℓ=ϕ(zℓ)(a0=x)z_\ell = W_\ell a_{\ell-1} + b_\ell,\quad a_\ell = \phi(z_\ell) \quad (a_0 = x)9 is a function solely of the network parameters and the sampled noise up to layer θ=(W,b)\theta = (W, b)0, while the residual θ=(W,b)\theta = (W, b)1 is uniformly bounded in norm. At the output layer,

θ=(W,b)\theta = (W, b)2

The prediction simplifies in the large-noise limit:

θ=(W,b)\theta = (W, b)3

Thus, as θ=(W,b)\theta = (W, b)4, predictions become independent of θ=(W,b)\theta = (W, b)5 and depend only on random features θ=(W,b)\theta = (W, b)6. This demonstrates that APEX suppresses input-specific signals and amplifies the structural, representation-level aspects embedded in the model.

Input perturbation, θ=(W,b)\theta = (W, b)7, is shown to be a constrained form of activation perturbation, as its induced change at layer θ=(W,b)\theta = (W, b)8 is

θ=(W,b)\theta = (W, b)9

where a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)0 is the Jacobian; input noise thus spans a low-dimensional subspace of the activation space, in contrast to the full-variance, unconstrained perturbations of APEX.

3. Probing Regimes and Interpretive Phenomena

There exists a qualitative dichotomy between small- and large-noise regimes:

  • Small-Noise Regime (a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)1): The residual a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)2 dominates. Predictions remain input-sensitive. In this regime, escape noise correlates strongly with sample regularity metrics such as memorization score and consistency/C-score (Spearman’s a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)3–a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)4 on ImageNet and CIFAR-100). APEX detects smooth semantic transitions in networks trained on controlled splits, exhibiting monotonic probability transfer that aligns with learned representations.
  • Large-Noise Regime (a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)5): The term a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)6 becomes dominant, rendering predictions input-agnostic. The network's output converges to a stationary, model-characteristic distribution. This regime exposes global biases: benign models exhibit high entropy output distributions, whereas backdoored models demonstrate collapse of output probability onto the target class (near-zero entropy).

The framework enables computation of normalized entropy,

a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)7

and quantification of target-class concentration, serving as diagnostics for backdoor detection and capacity-induced bias amplification.

4. Empirical Evaluations and Case Studies

APEX has been systematically validated through distinct probes:

Probe Type Quantitative Outcome Interpretation
Sample Regularity Spearman’s a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)8–a~ℓ(x;σ)=ϕ(zℓ(x))+σξℓ,ξℓ∼N(0,I)\tilde a_\ell(x; \sigma) = \phi(z_\ell(x)) + \sigma \xi_\ell, \quad \xi_\ell \sim \mathcal{N}(0, I)9 with memorization score Effective, lightweight alternative to ensembles
Random-Label Models Average escape noise decreases with more random labeling Reveals fragmented, non-semantic decision regions
Semantic Alignment Monotonic class transfer under activation noise only Confirms structure in representation space
Backdoor Detection Target class ℓ=1,…,L\ell = 1,\dots,L0–ℓ=1,…,L\ell = 1,\dots,L1 vs. ℓ=1,…,L\ell = 1,\dots,L2–ℓ=1,…,L\ell = 1,\dots,L3 (benign); entropy collapse Captures global, training-induced bias
Model Architecture Sensitivity Deeper ResNets: stronger probability collapse; ViTs: partial, attenuated collapse Architecture-dependent bias revelation

In all cases, input- or parameter-level perturbations fail to exhibit the monotonicity, transition alignment, or diagnostic sharpness furnished by APEX.

5. Computational Complexity and Practical Implementation

Each estimation for APEX requires ℓ=1,…,L\ell = 1,\dots,L4 full forward passes per input; ℓ=1,…,L\ell = 1,\dots,L5 is typical for CIFAR and ℓ=1,…,L\ell = 1,\dots,L6 for ImageNet. The noise injection itself constitutes a layerwise vector addition, which introduces minimal computational overhead. No model retraining or ensemble construction is necessary; all analysis occurs at inference with fixed weights and is trivially parallelizable over both examples and Monte Carlo samples. Choice of noise scale ℓ=1,…,L\ell = 1,\dots,L7 and threshold ℓ=1,…,L\ell = 1,\dots,L8 for metrics such as escape noise allows sensitivity–cost trade-offs.

APEX complements input and parameter perturbation probes—by acting directly on hidden representations, it accesses structural information that cannot be inferred from reachable input space alone.

6. Methodological Distinction and Conceptual Scope

APEX unifies local sample analysis and global model bias probing within a single, theoretically grounded framework. It admits input perturbation as a constrained, degenerate instance, subsuming prior approaches in expressive power. The ability to interpolate between regimes by modulating ℓ=1,…,L\ell = 1,\dots,L9 enables fine-grained scrutiny of memorization, regularity, semantic partitioning, and bias-induced collapse, revealing properties inaccessible to traditional probing techniques. The method's lightweight computational profile combined with its ability to interrogate internal network structure positions it as an effective probe for model interpretability, robustness diagnostics, and backdoor detection (Ren et al., 3 Feb 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Activation Perturbation for EXploration (APEX).