Papers
Topics
Authors
Recent
Search
2000 character limit reached

Cross-Mode Activation Patching in Neural Decoding

Updated 8 February 2026
  • Cross-mode activation patching is a causal interpretability technique that replaces internal neural activations across vocalized, mimed, and imagined speech modalities.
  • It employs coarse-to-fine tracing in convolutional and recurrent layers to identify compact subspaces that critically influence decoding performance.
  • Neuron-level interventions and manifold interpolation studies reveal that small, specific neuron subsets drive cross-modal transfer via graded latent representations.

Cross-mode activation patching is a causal mechanistic interpretability technique for probing neural networks trained on multimodal datasets, particularly in the context of brain-to-speech decoding. It systematically substitutes internal activations in a model across distinct input modalities—such as vocalized, mimed, and imagined speech—while holding all other activations and weights fixed. This framework enables researchers to localize, quantify, and characterize how representations supporting cross-modal generalization are encoded within neural architectures and to distinguish whether information is preserved in discrete, localizable subspaces or distributed activity patterns (Maghsoudi et al., 1 Feb 2026).

1. Formal Definition and Methodological Foundations

Let m,n{V,M,I}m,n\in\{\mathrm{V},\mathrm{M},\mathrm{I}\} denote modes corresponding to vocalized, mimed, and imagined speech, respectively. For a fixed decoder ff with LL layers, the activation at layer \ell for input xherox^{\text{hero}} in mode mm is a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}. The patching operator at layer \ell is formally defined as

P,mn(a(m))a(n),P_{\ell,m\rightarrow n}(a_\ell^{(m)}) \triangleq a_\ell^{(n)},

which replaces mode mm's activations by those from mode ff0 for paired linguistic content.

Given model split ff1, patched inference is

ff2

with all parameters fixed.

Causal impact is quantified by comparing the patched to unpatched output:

ff3

where ff4 is a metric such as negative Pearson correlation (PCC) or Mel Cepstral Distortion (MCD). Directionality is probed via patching from a higher- to lower-performing mode (sufficiency: ff5 or ff6) and in the reverse (necessity).

2. Causal Tracing: Identifying Localized Cross-Mode Structure

Coarse-to-fine tracing is employed to specify which internal subspaces are sufficient or necessary for cross-modal information transfer. In convolutional layers (ff7), the ff8 output channels are divided into four contiguous ff9-channel groups (LL0–LL1), each patched independently. Similarly, for a recurrent layer (LL2) of length LL3, time is segmented into thirds: Early LL4, Mid LL5, and Late LL6.

Coarse group patching (Table 4) finds that:

  • Sufficiency: Patching group LL7 from vocalized into imagined yields LL8 (from LL9), \ell0 (\ell1).
  • Necessity: Reverse patching \ell2 causes \ell3 to drop from \ell4 and \ell5 to rise \ell6.

Fine sliding-window tracing in the recurrent layer uses \ell7-timestep (25% of \ell8) windows, shifted in \ell9-step increments; windows spanning xherox^{\text{hero}}0 yield most of the cross-mode benefit, demonstrating local temporal specificity.

3. Subspace Identification and Causal Scrubbing

Coarse tracing reveals two localized regions mediating cross-mode transfer:

  • xherox^{\text{hero}}1: convolutional channels xherox^{\text{hero}}2 (16-dimensional)
  • xherox^{\text{hero}}3: xherox^{\text{hero}}4 recurrent time-steps xherox^{\text{hero}}5

Let xherox^{\text{hero}}6 be the selection mask for xherox^{\text{hero}}7, xherox^{\text{hero}}8 analogously for xherox^{\text{hero}}9. Subspace projections are defined as

mm0

Hybrid (“scrubbed”) activations are constructed by combining the subspace of interest from the donor mode with the complement from a random example in the target mode:

mm1

Variants include KEEP-Conv (only conv keep), KEEP-RNN, and KEEP-Combo. Performance is compared to random baselines (RAND-Conv, RAND-RNN) selecting the same number of features at random.

Key scrubbing sufficiency results (I←V direction):

Intervention PCC MCD
FullPatch 0.954 1.63
KEEP-Conv 0.666 3.13
RAND-Conv 0.564 3.23

A compact mm2-dimensional conv subspace suffices for a majority of the benefit, outperforming random selection and demonstrating directional asymmetry in transfer.

4. Neuron-Level Patching and Distributed Coding

Refinement proceeds to individual neuron and “top-k” group interventions. For source mm3, target mm4, patching neuron mm5 at all mm6 yields

mm7

The effect on decoding is measured by mm8.

For top-k patching, neurons are ranked by mean mm9; patching the top a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}0 produces a “saturation-degradation” curve: a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}1 rises, peaks at a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}2, then falls due to interference.

Key quantitative findings:

  • Peak a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}3 at a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}4 in RNN, a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}5 in Conv.
  • No single neuron produces a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}6; only a small, specific subset suffices.

Sentence-level coverage indicates that the top-5 RNN neurons, for the V→M case, suffice for up to a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}7 of sentences, with lower coverage in Conv, supporting the conclusion that cross-mode transfer is mediated by small neuron subsets, but not isolated units.

5. Manifold Structure: Tri-Modal Activation Interpolation

To interrogate the geometry of speech mode representations, activations at each layer are interpolated as

a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}8

For tri-modal generalization, convex combinations a(m)(xhero)Rda_\ell^{(m)}(x^{\text{hero}})\in\mathbb{R}^{d_\ell}9, where \ell0.

Observed effects include:

  • Nearly linear \ell1 and \ell2 transitions in \ell3 suggest graded local coding at mid-level layers.
  • RNN interpolation curves are smoother with mild non-linearities, indicating higher-level structure.
  • Mimed activations occupy an intermediate point between vocalized and imagined, not forming a discrete regime.

6. Quantitative Highlights and Interpretational Summary

Major experimental outcomes include:

  • Full layer activation patching: V→I (sufficiency) raises \ell4 from \ell5 to \ell6; I→V (necessity) drops \ell7 from \ell8 to \ell9.
  • A 16-channel conv block and a P,mn(a(m))a(n),P_{\ell,m\rightarrow n}(a_\ell^{(m)}) \triangleq a_\ell^{(n)},032-step RNN window account for nearly all patching benefit; random subspaces are significantly less effective.
  • Peak sufficiency is obtained by patching P,mn(a(m))a(n),P_{\ell,m\rightarrow n}(a_\ell^{(m)}) \triangleq a_\ell^{(n)},1 RNN or P,mn(a(m))a(n),P_{\ell,m\rightarrow n}(a_\ell^{(m)}) \triangleq a_\ell^{(n)},2 conv neurons, confirming the absence of isolated “magic” units.

All reported differences are statistically significant with P,mn(a(m))a(n),P_{\ell,m\rightarrow n}(a_\ell^{(m)}) \triangleq a_\ell^{(n)},3 for major conditions (P,mn(a(m))a(n),P_{\ell,m\rightarrow n}(a_\ell^{(m)}) \triangleq a_\ell^{(n)},4 sentences, 5-fold cross-validation). The results establish that speech modes lie on a single, graded manifold in latent space, and cross-mode transfer is driven by compact, layer-specific subspaces rather than widely distributed or individual-neuron activity. Directionality is pronounced: only vocalized representations enable effective transfer to imagined/mimed decoding; the reverse patching severely degrades performance (Maghsoudi et al., 1 Feb 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Cross-Mode Activation Patching.