Papers
Topics
Authors
Recent
Search
2000 character limit reached

Patterning by Susceptibility Inversion

Updated 13 July 2026
  • The paper introduces a method that uses Bayesian learning and linear response theory to compute minimal data perturbations that induce desired changes in a model’s internal structure.
  • It employs susceptibility matrices to measure how posterior expectations shift under data reweighting, with the Moore–Penrose pseudoinverse mapping target shifts to optimal perturbations.
  • Empirical examples in language models and synthetic tasks demonstrate that structurally informed data interventions can steer internal computations, albeit with scaling and estimation challenges.

Patterning by susceptibility inversion is the use of linear response theory in Bayesian learning to determine which infinitesimal perturbations of the training data distribution produce a desired change in a model’s internal structure. In this framework, internal structure is represented by observables on parameter space, susceptibilities quantify how posterior expectation values of those observables respond to data reweighting, and the inverse problem—called patterning—is solved by applying the Moore–Penrose pseudoinverse to the susceptibility matrix. The formulation developed in "Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning" and "Patterning: The Dual of Interpretability" places the method within a statistical-mechanical view of training, where the same mathematical object used to read structure can be inverted to write it (Elliott et al., 8 May 2026, Wang et al., 20 Jan 2026).

1. Conceptual position within interpretability

Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. Patterning is introduced as the dual problem: given a desired form of generalization, determine what training data produces it (Wang et al., 20 Jan 2026). In this duality, interpretability is associated with extracting and analyzing latent structure, whereas patterning treats those same structural quantities as targets for intervention.

The central quantities are observables

ϕi(w):WR\phi_i(w):W\to\mathbb R

defined on parameter space. These observables may isolate the contribution of an attention head, a layer, or a local learning-coefficient estimator. Their posterior expectations serve as structural coordinates. Patterning then asks for a perturbation of the data distribution that shifts these coordinates by a prescribed amount.

A key feature of this formulation is that it is explicitly Bayesian. In modern Bayesian deep learning, a trained network is represented by a Gibbs posterior

Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,

where L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)], nn is the effective sample size, and β>0\beta>0 is an inverse-temperature or tempering parameter (Elliott et al., 8 May 2026). The corresponding empirical construction in the finite-sample setting is the annealed posterior

pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).

Within this view, patterning is not a heuristic for dataset editing; it is the inversion of a local differential map from data distributions to posterior expectation values.

This suggests that patterning is structurally tied to the geometry of the posterior rather than to purely input-output behavior. A plausible implication is that the method is most naturally interpreted as an intervention on learned computation, not merely on surface performance.

2. Susceptibility as linear response

Susceptibilities quantify how posterior expectations shift when the data distribution is perturbed. If a probe distribution qq' is introduced through

qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],

and ϕh\langle \phi\rangle_h denotes the posterior expectation of an observable under the perturbed loss, then the susceptibility of ϕ\phi toward the perturbation Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,0 is

Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,1

The fluctuation–dissipation theorem identifies this derivative with a posterior covariance: Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,2 where Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,3 (Elliott et al., 8 May 2026).

For a point perturbation, one obtains the per-sample susceptibility. With Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,4,

Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,5

In the notation of the patterning paper, for a sample Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,6,

Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,7

with Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,8 and Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,9 (Wang et al., 20 Jan 2026).

This covariance formula is the core analytic step. It converts a derivative with respect to a data-distribution perturbation into a posterior statistic. The theory therefore links sensitivity analysis, Bayesian influence, and statistical mechanics within a single local-response object. The primer further notes that per-sample losses yield the influence matrix, described there as the Bayesian influence function of (Kreer et al., 30 Sep 2025), while component-localized observables yield the structural susceptibility matrix (Elliott et al., 8 May 2026).

A common misunderstanding would be to treat susceptibility as a global descriptor of training dynamics. The formulation is explicitly first-order: it describes the initial response of posterior expectations to an infinitesimal perturbation.

3. Structural susceptibility matrices and data-to-structure Jacobians

When a family of observables L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]0 is chosen, often to localize to particular model components, and a basis of perturbation directions L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]1 is fixed, the pairwise susceptibilities assemble into the structural susceptibility matrix

L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]2

Equivalently, if local perturbations of the data distribution are parameterized by coordinates L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]3, then

L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]4

If the structural coordinate map is written as

L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]5

its Jacobian at L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]6 is exactly L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]7: L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]8 The primer states this explicitly: the susceptibility matrix is, up to a factor of L(w)=Ezq[z(w)]L(w)=\mathbb E_{z\sim q}[\ell_z(w)]9, the Jacobian of the map from data distributions to structural coordinates (Elliott et al., 8 May 2026).

The corresponding formulation in the patterning paper writes the same relation in integral and coordinate form. For an infinitesimal density shift nn0,

nn1

and in coordinates,

nn2

This matrix viewpoint has two consequences. First, it turns interpretability into a differential map from sample-weight space to observable space. Second, it makes singular-value analysis natural: the left singular vectors identify principal structures in observable space, and the right singular vectors identify principal data patterns in sample-weight space (Wang et al., 20 Jan 2026).

In practical terms, nn3 or nn4 is the local coupling matrix between data patterns and model components. This suggests that the matrix provides a structured decomposition of which data perturbations are most effective at modulating specific internal computations.

4. Pseudoinverse solutions to the patterning problem

Patterning is defined as the inverse problem: given a desired infinitesimal change

nn5

or, equivalently, a target shift nn6 in structural coordinates, determine the smallest-norm perturbation in data space that realizes it to first order. In the notation of the primer, one solves

nn7

Because nn8 is generally non-square, the solution is posed as a minimum-Euclidean-norm problem. Introducing Lagrange multipliers yields

nn9

where

β>0\beta>00

is the Moore–Penrose pseudo-inverse of β>0\beta>01 on the row-space of β>0\beta>02. Thus the minimal-norm patterning perturbation is

β>0\beta>03

(Elliott et al., 8 May 2026).

The patterning paper presents the same derivation in terms of the susceptibility matrix β>0\beta>04: β>0\beta>05 with solution

β>0\beta>06

If regularization is desired, one may instead solve

β>0\beta>07

whose closed-form solution is

β>0\beta>08

(Wang et al., 20 Jan 2026).

Singular-value decomposition gives an especially transparent interpretation. Writing

β>0\beta>09

with singular values pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).0, the vectors pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).1 are principal structures and pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).2 are principal data patterns. The pseudoinverse becomes

pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).3

When the target is aligned with a single mode, pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).4, the minimal intervention is

pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).5

This is the algebraic core of susceptibility inversion: structural targets are projected onto accessible response modes, and the corresponding data intervention is read off from the right singular vectors (Wang et al., 20 Jan 2026).

5. Empirical estimation and operational approximations

The formal theory is stated at the population level, but practical patterning uses an empirical posterior

pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).6

constructed from a finite training set pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).7. Population covariances are then replaced by the empirical estimator

pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).8

The primer states that this estimator is exactly the per-sample weight derivative, identified there with Gustafson’s local-sensitivity or the Bayesian influence function, of pnβ(w)=1Znβexp{nβLn(w)}φ(w).p_n^\beta(w)=\frac{1}{Z_n^\beta}\,\exp\{-n\beta\,L_n(w)\}\,\varphi(w).9 when sample qq'0 is upweighted by an infinitesimal amount (Elliott et al., 8 May 2026).

For component-localized observables qq'1, the complementary parameters must be clamped and only the component subspace sampled, a procedure described as weight-restriction. This induces a multiplicative renormalization factor. In practice, that constant is removed by standardizing each row or column, for example by subtracting its mean and dividing by its standard deviation across data points. Posterior expectations and covariances are approximated by SGLD or Langevin-dynamics sampling, typically on minibatches, which adds an extra sampling noise layer. These approximations define an empirical susceptibility matrix qq'2, and its pseudo-inverse is often ridge-regularized: qq'3 Finite reweightings of the training batch are then computed as

qq'4

(Elliott et al., 8 May 2026).

The patterning paper provides the experimental sampling configurations used in two benchmark settings. In the induction-circuit experiments, susceptibilities were estimated by SGLD sampling with qq'5, qq'6, step size qq'7, qq'8 chains, and qq'9 draws. In the parentheses-balancing experiments, susceptibilities were measured on qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],0 in-distribution samples with SGLD using qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],1, qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],2, step size qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],3, qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],4 chains, and qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],5 draws (Wang et al., 20 Jan 2026).

The framework therefore depends on two layers of approximation: a linear-response approximation in data space and a sampling-based approximation in posterior space. A common misconception is that pseudoinverse patterning directly prescribes finite dataset edits without caveats. The underlying papers instead emphasize small perturbations, estimator noise, and regularization.

6. Experimental demonstrations and scope

Two demonstrations are central in the published exposition: an induction-circuit experiment in a small LLM and a synthetic parentheses-balancing task. Together they instantiate both principal-direction patterning and susceptibility-gap patterning.

Setting Intervention basis Reported outcome
Induction circuit Re-weighting by qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],6 Accelerates or delays formation of structure
Parentheses balancing Susceptibility gap qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],7 Steers algorithmic preference

In the small LLM, the model is a 2-layer, 8-head GPT-2–style transformer with qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],8 M parameters, trained for qh=(1h)q+hq,h[0,1],q_h=(1-h)\,q+h\,q',\qquad h\in[0,1],9 k steps on a Pile subset. The SVD of the ϕh\langle \phi\rangle_h0 (≈ ϕh\langle \phi\rangle_h1 M-token) susceptibility matrix yields the second singular mode ϕh\langle \phi\rangle_h2. This mode was found to couple the well-known induction circuit in weights to induction patterns in the data. Data interventions were defined as four per-token weighting maps based on the scalar ϕh\langle \phi\rangle_h3 loading: Suppress, Baseline, Induce2×, and Induce4×. Retraining used ϕh\langle \phi\rangle_h4 seeds each and measurements every ϕh\langle \phi\rangle_h5 k steps. The reported metrics were prefix-matching score, previous-token attention score, and UMAP principal component explained variance. The reported result is that repressing tokens with negative ϕh\langle \phi\rangle_h6 nearly prevents induction-circuit formation, while amplifying them ϕh\langle \phi\rangle_h7 or ϕh\langle \phi\rangle_h8 accelerates and strengthens it (Wang et al., 20 Jan 2026).

The primer summarizes the same experiment by stating that the second principal susceptibility mode ϕh\langle \phi\rangle_h9 identifies a coherent token-pattern that elicits an induction-like loop in the network, and that reweighting each token’s loss according to its loading on ϕ\phi0, using a small multiplier ϕ\phi1, can steer the model toward or away from producing induction behaviors (Elliott et al., 8 May 2026).

In the synthetic parentheses-balancing task, there are two distinct zero-loss solutions on the training distribution: #Nested, which accepts only well-nested sequences, and #Equal-Count, which accepts any sequence with equal numbers of “(” and “)”. Singular learning theory associates to each local minimum ϕ\phi2 a local learning coefficient

ϕ\phi3

and among equal-loss solutions, posterior odds favor the one with lower LLC. Two observables, ϕ\phi4 and ϕ\phi5, isolate the loss around each solution, with ϕ\phi6 and ϕ\phi7. Measuring ϕ\phi8 and ϕ\phi9 reveals how each sample shifts these coefficients. The susceptibility gap

Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,00

then identifies the influential examples for steering between the two solution classes (Wang et al., 20 Jan 2026).

The experimental protocol used Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,01 transformers from Li et al. (2025), each with Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,02–Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,03 layers, Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,04 heads, and weight decay Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,05, exhibiting OOD accuracies clustered near Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,06 or Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,07. The top-Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,08 Nested and bottom-Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,09 Equal-Count models were identified, and their susceptibilities averaged to approximate the two rows of Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,10. “Almost nested” and “almost equal” samples were then selected by maximizing or minimizing the susceptibility gap. Two Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,11 k-sample datasets were constructed: Almost Nested, which removed Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,12 k False samples and added two copies of Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,13 k “almost nested”; and Almost Equal, which removed Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,14 k False and added four copies of Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,15 k “almost equal.” Retraining Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,16 models on each dataset produced the reported result that Almost Nested eliminates high-OOD solutions, with all Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,17 converging to Equal-Count, while Almost Equal increases the fraction of Nested solutions, with mean OOD Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,18 versus Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,19 baseline (Wang et al., 20 Jan 2026).

These experiments are also notable because the direction of the effect is not reducible to simple data augmentation intuitions. In the parentheses task, the intervention acts through the susceptibility gap and the local learning coefficients rather than through an obvious direct encoding of the target rule. A plausible implication is that patterning can select among competing zero-loss mechanisms by reshaping posterior preference over solution basins.

7. Limitations, scaling questions, and relation to adjacent methods

The framework has explicitly stated limits. All of the formal derivations rest on a linear-response approximation, valid when the perturbation is small and when the map from perturbation parameters to structural coordinates is well approximated by its first derivative. The scaling by Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,20 controls the sharpness of the posterior, and the inversion can amplify estimation errors in small singular-value directions, which is why ridge regularization is used in practice (Elliott et al., 8 May 2026).

The current experiments are also bounded in scale and protocol. The patterning paper states that the present demonstrations use approximately Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,21 M-parameter models and offline one-off interventions. It further states that susceptibility estimation via SGLD is computationally expensive, approximately Π(dw)exp(nβL(w))π(w)dw,\Pi(dw)\propto \exp\bigl(-n\beta\,L(w)\bigr)\,\pi(w)\,dw,22 a single training run for the induction experiments, while also noting that it scales linearly with model size and may be accelerated by approximate predictors or online protocols (Wang et al., 20 Jan 2026).

Several extensions are identified in the source material. These include online patterning, in which data weights are adjusted dynamically as structures emerge; scaling by sampling only a subset of components or using learned approximators for susceptibilities; and alignment-oriented uses, in which patterning constrains or encourages specific internal computations, such as preventing deceptive structures or promoting robust constraint-encoding (Wang et al., 20 Jan 2026). These are presented as extensions rather than established empirical conclusions.

Patterning by susceptibility inversion is therefore best understood as a local, posterior-geometric method for shaping internal structure. It does not replace mechanistic interpretability; rather, it formalizes the inverse operation. The same susceptibility matrix that reads structure through covariance and singular modes becomes, through pseudoinverse inversion, a prescription for minimal data reweighting that writes structure back into the model (Elliott et al., 8 May 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Patterning by Susceptibility Inversion.