---
title: Cross-Variation Patching in Transformers
url: https://www.emergentmind.com/topics/cross-variation-patching
type: topic
---

# Cross-Variation Patching in Transformers

Searching arXiv for the primary paper and closely related work on patching diagnostics and cross-patching.
Cross-variation patching, in the continuous-depth field-theoretic language of Olivieri & Pérez Rodríguez, treats a Transformer’s residual stream as a field over depth and token position and transfers the residual-field difference between two prompt variants by inserting that difference as a localized source at a chosen site. In this formulation, patching is not only an intervention but also a prediction problem: the downstream effect of the injected variation can be estimated from a first-order sensitivity field or from an empirical Green-function slice, and the same framework supports patch-site inference and cross-scale transfer [2605.25225]. Related work uses cross-patching to separate upstream state from late readout across pretrained and instruction-tuned checkpoints, showing that a late-layer effect need not be self-contained [2605.07284].

## 1. Continuous-depth residual-field formulation

The starting point is a continuous “depth” coordinate $t \in [0,T]$ with $t_\ell=\ell\,T/L$ and a token-space coordinate $x$. The residual stream of a pre-norm Transformer is represented as a field
$$
r(t,x)\in\mathbb R^{d_{\rm model}},\qquad t=0,\dots,T,\;x=1,\dots,n.
$$
In the absence of interventions, the field evolves by the residual ODE
$$
\partial_t\,r(t,x)=F_t[r](x),
$$
where $F_t$ encodes attention $+$ MLP updates at depth $t$ [2605.25225].

A discrete activation patch at layer $t_0$ and token $x_0$ is modeled as an impulsive source term $\delta J(t,x)$:
$$
\partial_t\,r(t,x)=F_t[r](x)+\delta J(t,x),\qquad
\delta J(t,x)=\delta(t-t_0)\,\delta_{x,x_0}\,\Delta r_0,
$$
with
$$
\Delta r_0=r^{\rm src}(t_0,x_0)-r^{\rm base}(t_0,x_0).
$$
This makes patching a localized source insertion rather than an ad hoc replacement rule. Within this representation, a patch is specified by its support in depth-token space and by the vector it injects.

The significance of this reformulation is organizational as well as mathematical. The abstract explicitly states that the framework treats “patching as localized source insertion,” “patch effects as sensitivity-field predictions,” and “downstream propagation as empirical Green-function response,” thereby providing a common language for activation patching, causal tracing, path patching, and steering directions [2605.25225].

## 2. Sensitivity fields, adjoints, and empirical Green functions

For a scalar output $y$, such as a logit difference at the final token, a small source $\delta J$ induces the first-order change
$$
\delta y
=\int_0^T dt\int dx\;a(t,x)\cdot\delta J(t,x)+O(\|\delta J\|^2),
$$
where the adjoint or “sensitivity field” is
$$
a(t,x)=\frac{\delta y}{\delta r(t,x)}=S_{\rm out}(t,x)\in\mathbb R^{d_{\rm model}}.
$$
For a small source inserted at a single site $(t_0,x_0)$,
$$
\delta y\approx a(t_0,x_0)\cdot \delta J(t_0,x_0).
$$
The quantity $a(t,x)$ is therefore the local first-order predictor of patch efficacy [2605.25225].

The sensitivity field is obtained from an adjoint construction. Introducing an adjoint field $\lambda(t,x)$ and the action
$$
S[r,\lambda]
=\int_0^T dt\int dx\;\lambda\bigl(\partial_t r - F_t[r]\bigr)
+ S_{\rm readout}[r(T,\cdot)],
$$
stationarity $\delta S=0$ gives the backward equation
$$
-\,\partial_t\,\lambda(t,x)
=\int dy\;L_t^\top(y,x)\,\lambda(t,y),\qquad
\lambda(T,x)=\frac{\partial y}{\partial r(T,x)},
$$
whose solution is $\lambda(t,x)=a(t,x)$. In this sense, the backward pass defines the same response object that first-order patch prediction uses.

Linearizing around the unpatched trajectory $r^0$ by writing $r=r^0+\delta r$ yields
$$
\partial_t\,\delta r(t,x)
=\int dy\;L_t(x,y)\,\delta r(t,y)+\delta J(t,x),
$$
with $L_t=\delta F_t/\delta r\vert_{r^0}$. The corresponding fundamental solution is the Green’s function
$$
G(t,x;\,t_0,x_0)=\frac{\delta\,r(t,x)}{\delta J(t_0,x_0)},
$$
satisfying
$$
\partial_t\,G(t,\cdot;t_0,x_0)=L_t\,G(t,\cdot;t_0,x_0),\qquad
G(t_0,x;t_0,x_0)=\delta_{x,x_0}\,I_{d\times d}.
$$
Once $G$ is known, any localized patch propagates by
$$
\delta r(t,x)=\int dt_0\,dx_0\;G(t,x;\,t_0,x_0)\,\delta J(t_0,x_0).
$$
The paper’s abstract reports empirical measurement of “structured anisotropic propagation across depth and token position” and construction of “response descriptions from high-sensitivity sites and sliced Green operators,” which situates these objects as experimentally measurable rather than purely formal [2605.25225].

## 3. Transfer between prompt variants

The canonical cross-variation setting considers two prompt variants $q_A$ and $q_B$ with residual fields $r_A$ and $r_B$. At a fixed site $(t_0,x_0)$, the patch-direction is
$$
J_{A\to B}=r_B(t_0,x_0)-r_A(t_0,x_0).
$$
Its first-order scalar effect on run $A$ is predicted by
$$
\delta y\approx a_A(t_0,x_0)\cdot J_{A\to B}.
$$
Operationally, one computes $a_A(t_0,x_0)$ by backpass on the $q_A$ run and measures the inner product with $J_{A\to B}$. If it is large and positive, one expects the patched run to exhibit behavior closer to $q_B$ [2605.25225].

A more precise transfer uses a Green slice. For a downstream readout at $(t_{\rm out},x_{\rm out})$,
$$
\delta r(t_{\rm out},x_{\rm out})
=G(t_{\rm out},x_{\rm out};\,t_0,x_0)\;J_{A\to B}.
$$
This supports site choice by maximizing $\|G\cdot J_{A\to B}\|$ over candidate locations. The field-theoretic description therefore distinguishes two predictive modes: a scalar first-order criterion using the sensitivity field and a downstream-state criterion using the empirical Green operator.

The worked toy example uses $q_A=$ “The capital of Spain is” and $q_B=$ “The capital of Italy is,” with the model predicting “ Madrid” and “ Rome,” respectively. Both runs are taken to depth $t_0=10$ at final token $x_0=-1$, the patch-direction
$$
J_{A\to B}=r_B(10,-1)-r_A(10,-1)
$$
is recorded, and backprop on $q_A$ yields
$$
a_A(10,-1)=\partial\,(\text{logit}_{\text{“Rome”}}-\text{logit}_{\text{“Madrid”}})/\partial r_A(10,-1).
$$
The predicted effect is
$$
\delta y_{\rm pred}=a_A(10,-1)\cdot J_{A\to B}.
$$
If $\delta y_{\rm pred}\gg 0$, the patch is expected to favor “Rome” over “Madrid”; after applying
$$
r_A(10,-1)\longmapsto r_A(10,-1)+J_{A\to B},
$$
the “Rome” logit rises, often overtaking “Madrid,” thereby effecting a cross-variation transfer of the capital-fact behavior [2605.25225].

## 4. Patch-site inference and the local linear regime

The same formalism yields an adjoint variational problem for optimal patch-site selection. To maximize a desired $\Delta$-behavior $\Delta y^\star$, one may solve
$$
\min_{J(\cdot)}\;C\bigl[J\bigr]
\quad\text{s.t.}\quad
\int dt\,dx\;a(t,x)\cdot J(t,x)=\Delta y^\star,
\quad
\partial_t r=F_t[r]+J.
$$
In the linear regime, where $\delta y=a\cdot J$, and with the energy penalty
$$
C[J]=\tfrac12\!\int\|J\|^2,
$$
the Lagrange-multiplier condition gives
$$
J^\star(t,x)\propto a(t,x).
$$
Thus the optimal source aligns with the sensitivity field, while a sparsity or site-support constraint restricts the support of $J$ to a small set of $(t,x)$ and thereby selects a handful of patch sites [2605.25225].

This formulation places site selection on the same footing as forward-response prediction. Instead of searching over interventions solely by brute-force patching, one first computes response objects and then uses them to propose high-value sites. The abstract reports that the paper identifies “a bounded local linear regime” and predicts “patch effects from first-order sensitivities across residual sites,” which is the empirical condition under which the proportionality $J^\star\propto a$ is useful in practice [2605.25225].

A common misconception is to treat any successful patch as intrinsically local and self-explanatory. The field-theoretic treatment suggests a stricter interpretation: a local intervention is only one element of a distributed response, and its effect depends on both its first-order sensitivity and its downstream propagation. This suggests that causal claims from patching are strongest when accompanied by response objects—sensitivities, propagated fields, or Green-operator slices—rather than by endpoint behavior alone.

## 5. Checkpoint recombination and first-divergence cross-patching

A distinct but closely related diagnostic is “first-divergence cross-patching,” introduced to study cooperation between earlier computation and the late stack in pretrained base (PT) and instruction-tuned (IT) checkpoints. Let $\ell^\*$ be the depth at which the “late-stack boundary” is drawn, typically $\approx 60\%$ of total depth. At the first token where PT and IT disagree under greedy sampling, the shared history is $p^\*$ and the divergent tokens are
$$
t_{PT}=\arg\max \logit_{PT}(\cdot\mid p^\*),\qquad
t_{IT}=\arg\max \logit_{IT}(\cdot\mid p^\*).
$$
The protocol constructs four hybrid forward passes by mixing upstream state from one checkpoint with the late stack from either checkpoint:
$$
z(UP,LT)=f_{>\ell^\*}^{LT}\!\bigl(h^{\ell^\*}_{UP}(p^\*)\bigr),
$$
and studies the divergent-token margin
$$
Y(UP,LT)=z(UP,LT)_{t_{IT}}-z(UP,LT)_{t_{PT}}.
$$
This yields
$$
\Delta_{PT\_up}=Y(PT,IT)-Y(PT,PT),
\qquad
\Delta_{IT\_up}=Y(IT,IT)-Y(IT,PT),
$$
and the interaction
$$
\text{Interaction}=\Delta_{IT\_up}-\Delta_{PT\_up}.
$$
Across the Core-5 dense families (4B–32B), the reported point estimates are $\Delta_{PT\_up}\approx +0.76$, $\Delta_{IT\_up}\approx +2.44$, and $\text{Interaction}\approx +1.68$ logits; the interaction is positive in every family [2605.07284].

The interpretation given in the paper is that the IT late stack has a real PT-upstream effect, but its larger effect in the IT checkpoint appears only when it reads its own post-trained upstream state. The reported “Portable (PT-upstream) share” is $\Delta_{PT\_up}/\Delta_{IT\_up}\approx 31\%$ on average, with family range $19\%$–$44\%$, leaving the remaining $\approx 69\%$ of the IT late-stack effect dependent on IT upstream state. Sparse final-MLP features partially mediate this interaction: ablating the top-200 features in the IT late stack reduces the interaction by $26$–$48\%$, patching those features back into the weak hybrid rescues $\approx 0.5$ logits, and an earlier-layer patch into the final stack recovers $+1.71$ logits, with $\approx 0.13$ logit mediated by those same final-layer features [2605.07284].

The paper also reports a structured boundary-state closure result for Llama: injecting a rank-256 approximation of the descendant-minus-base $\ell^\*$-level residual shift into the weak hybrid recovers $71\%$ of the missing margin, while the full-delta recovers $\approx 97\%$; random or sign-flipped directions have near-zero or negative effect. Forced-token scoring further shows that the local token choice can change later exact-answer success: on CONTENT-REASON exact-answer prompts, forcing $t_{IT}$ yields a $+0.157$ gain in suffix-only exact-match success, with smaller positive effects on safety ($\approx +0.039$) and format ($\approx +0.026$) validators [2605.07284].

The main caution is explicit: when a behavior is localized to late layers, the late-stack effect should be tested under the other checkpoint’s upstream state before being treated as self-contained. In the paper’s phrasing, first-divergence cross-patching separates a direct late-stack component from an upstream-dependent component, and in the reported experiments most of the IT late-stack effect depends on the model’s own upstream computations [2605.07284].

## 6. Related notions in hardware patchability and software repair

Outside mechanistic interpretability, cross-variation comparison appears in hardware patch-design analysis. “Theoretical Patchability Quantification for IP-Level Hardware Patching Designs” defines patchability as a combination of controllability and observability and uses this to compare IP variations or patching-logic choices at RTL. For a signal $v$, the overall metric is
$$
P_v=w_C\cdot C_v+w_O\cdot O_v,
$$
with the paper taking all weights to be $1$, so $P_v=C_v+O_v$. A cross-variation comparison parses each RTL variant, marks directly patched nets, propagates $C$ and $O$ through the dataflow graph, computes $P_{total}$ or $P_{avg}$, and compares patchability under equal investment budgets. In the reported case study on “reglk_wrapper,” configuration V3 invests $110$ bits and yields $P_{avg}\approx 0.85$, whereas V4 invests more ($192$ bits) but gets only $P_{avg}\approx 0.54$, so V3 is strictly better [2311.03818].

A different software-repair use of cross-variation operates over candidate patches rather than activations. “A Single Patch Is Not Enough: Deterministic Fusion of Repair Candidates” defines a pool $P$ of candidate patches, decomposes each patch into edit atoms $\alpha=(I,R)$, measures pairwise similarity by Jaccard over edit-atom sets,
$$
J(p_i,p_j)=\frac{|E(p_i)\cap E(p_j)|}{|E(p_i)\cup E(p_j)|},
$$
builds repair neighborhoods from agreement graphs, selects a representative by medoid-style agreement, and applies evidence-constrained fusion (ECF) to retain repeated edit atoms and prune unsupported parts. On PatchFuseBench, PatchFusion solves $426/500$ bugs on SWE-bench Verified, $236/300$ on SWE-bench Multilingual, and reaches $87/371$ plausible patches on Defects4J; ECF alone adds $+5/+6/+9$ solved with zero regressions [2607.01597]. This is not activation patching, but it is a related instance of using structured variation across candidates to construct a stronger patch.

Multi-hunk repair studies formalize variation within a single patch. “Characterizing Multi-Hunk Patches: Divergence, Proximity, and LLM Repair Challenges” defines overall hunk divergence
$$
\mathrm{Div}(P)
=\ln(n)\times\frac{2}{n(n-1)}\sum_{1\le i<j\le n}\mathrm{Div}(h_i,h_j),
$$
together with a five-class spatial proximity taxonomy: Nucleus, Cluster, Orbit, Sprawl, and Fragment. On Hunk4J, which consists of $372$ real-world multi-hunk bugs mined from Defects4J, model success rates decline with increased divergence and spatial dispersion; with vanilla prompts, Plausible@1 rates are $12$–$27\%$, and no model succeeds in the most dispersed Fragment class [2506.04418]. A plausible implication is that cross-variation patching becomes more difficult as the variation to be transferred or coordinated becomes more heterogeneous and more spatially dispersed.

Taken together, these adjacent literatures reinforce a common methodological point. Whether the object being patched is a residual stream, a checkpoint boundary state, an RTL design, or a candidate repair pool, cross-variation methods are most informative when they explicitly model what varies, where it is inserted or compared, and how the induced effect propagates.

Source: https://www.emergentmind.com/topics/cross-variation-patching