---
title: Causal Input Normalization
url: https://www.emergentmind.com/topics/causal-input-normalization
type: topic
---

# Causal Input Normalization

Searching arXiv for recent and directly relevant papers on causal input normalization, causal normalizing flows, and counterfactual normalization.
Causal input normalization denotes a family of procedures that normalize inputs by reference to causal structure rather than by moment matching alone. In the cited literature, this idea appears in several technically distinct forms: counterfactual harmonization of observed features with structural causal models and normalizing flows, transformation of observations into exogenous variables through causally structured flows, normalization of time-series causality by a recipient-specific entropy budget, counterfactual removal of unstable causal pathways under dataset shift, and causal recurrent modulation of feature statistics in streaming systems [2106.06845] [2306.05415] [1501.03548] [1808.03253] [2605.25308].

## 1. Conceptual foundations

A recurrent theme is the decomposition of an observed input into a part attributable to explicit causes and a part attributable to exogenous variation. In structural causal model notation, this takes the form
$$
X_i = f_i(X_{\mathrm{pa}(i)}, U_i),
$$
with endogenous variables $X_i$, parents $X_{\mathrm{pa}(i)}$, and mutually independent exogenous variables $U_i$. In flow-based formulations, the inverse map from observations to exogenous variables functions as a causal normalization map: it removes the effect of parents and expresses each variable in terms of a noise coordinate that is independent under the base distribution [2306.05415].

A plausible unifying view is that causal input normalization replaces purely associational standardization with one of three causal operations. The first is **counterfactual standardization**, in which an observed input is regenerated under a fixed causal regime, such as a reference site or scanner. The second is **exogenous-coordinate normalization**, in which an invertible model maps data to independent noise variables while preserving parent-child structure. The third is **causal-share normalization**, in which the importance of a causal input is expressed relative to the total uncertainty budget of the recipient variable [2106.06845] [2412.12401] [1501.03548].

The literature also differs on what is being normalized. In some work, the normalized object is the observed feature vector itself, produced as a counterfactual. In some work, it is the latent variable $U$ recovered by an autoregressive or masked flow. In time-series causality, it is neither the observation nor the latent state, but the magnitude of causal influence, normalized relative to self-dynamics and noise. In dataset-shift settings, normalization refers to replacing factual inputs by counterfactual variables that exclude unstable parental effects [2011.02268] [2104.11360] [1808.03253].

## 2. Counterfactual harmonization of observed inputs

In flow-based harmonization of medical data, brain MRI ROI volumes are modeled with an SCM whose main observed variables are $x$ (145-dimensional ROI volumes), sex $s$, age $a$, and site $t$, together with exogenous noises $\epsilon_x,\epsilon_s,\epsilon_a,\epsilon_t$. The causal parents of $x$ are $s,a,t$ and $\epsilon_x$, with structural equations
$$
s := \epsilon_S,\qquad
a := f_A(\epsilon_A),\qquad
t := \epsilon_T,\qquad
x := f_X(\epsilon_X; s,a,t),
$$
where $f_X$ is implemented by a conditional normalizing flow. Training is by maximum likelihood with factorization
$$
p_\theta(x,s,a,t)=p_\theta(x\mid s,a,t)\,p_\theta(a)\,p_\theta(s)\,p_\theta(t).
$$
Harmonization is the abduction–action–prediction pipeline: infer $\epsilon_X=f_{X,\theta}^{-1}(x;s,a,t)$, intervene with $\mathrm{do}(t=\tau)$, and regenerate
$$
x^{\mathrm{harm}}=f_{X,\theta}(\epsilon_X; s,a,t=\tau).
$$
This makes the normalized input a counterfactual ROI vector under a standardized site, while preserving subject-specific variability in $\epsilon_X$ [2106.06845].

The practical significance of this formulation is that harmonization is performed in the original feature space rather than only in a latent space. The method distinguishes site effects from exogenous subject-specific factors, and does so by intervention on causal parents rather than by aligning marginal means and variances. The paper explicitly contrasts this with ComBat and related methods, arguing that mean-variance alignment can remove biologically meaningful variance, whereas the flow-based SCM preserves unknown confounders and subject-specific variability in $\epsilon_X$ [2106.06845].

Empirically, the method was evaluated on the iSTAGING consortium for age regression and on ADNI for Alzheimer’s disease classification. With BLSA-3T as source, the Flow-based SCM Quadratic Spline achieved MAE 6.92 on BLSA-1.5T, 6.44 on UKBB, and 15.68 on SHIP, compared with SrcOnly values 7.21, 7.27, and 17.14. For AD classification, Flow-based SCM Q-Spline reached 73.7% for ADNI-1 $\rightarrow$ ADNI-2 and 73.3% for ADNI-2 $\rightarrow$ ADNI-1, compared with 71.9% and 70.4% for SrcOnly. These results are presented as improved cross-domain generalization from causally harmonized inputs [2106.06845].

## 3. Flow-based causal normalization and causal consistency

Autoregressive normalizing flows furnish a second major interpretation of causal input normalization: the flow itself is treated as the SCM. In the bivariate CAREFL construction, an affine autoregressive flow with fixed ordering $\pi$ yields structural equations of the form
$$
x_j=e^{s_j(\mathbf{x}_{<\pi(j)})}z_j+t_j(\mathbf{x}_{<\pi(j)}),
$$
with statistically independent $z_j$. Under the paper’s assumptions, the permutation $\pi$ is interpreted as a causal ordering, the latent $z_j$ are exogenous noises, and the model is identifiable in the bivariate case when the linking function is non-linear and invertible. This produces a direct notion of causal normalization: the inverse flow maps observations to independent exogenous coordinates aligned with a causal ordering, and interventions or counterfactuals are implemented by modifying the relevant latent coordinate and pushing it through the flow [2011.02268].

Subsequent work on causal normalizing flows generalizes this view by using autoregressive masks determined by a known ordering or graph, so that $\flow:x\mapsto u$ is both a density model and a causal encoder. Under the non-linear ICA argument used there, if the learned flow matches the observational distribution of the SCM, then the latent variables are identifiable up to component-wise invertible transformations. The same work states the causal-consistency pattern through Jacobians,
$$
\nabla_x \flow(x)\equiv I-\mathcal{A},\qquad
\nabla_u \flow^{-1}(u)\equiv I+\sum_{n=1}^{\mathrm{diam}(\mathcal{A})}\mathcal{A}^n,
$$
in the sense of structural equivalence. Its recommended practical design is a single-layer abductive model with full graph information, so that the map from observations to exogenous variables functions as a structured normalization aligned with parent sets [2306.05415].

A more recent line identifies **causal inconsistency** as a mismatch between the graph encoded by a normalizing flow and the graph specified by an SCM. The Causally Consistent Normalizing Flow constructs a sequential representation of the SCM through topological batching,
$$
\mathbf{B}=(\mathbf{B}_1,\dots,\mathbf{B}_n),
$$
and composes partial causal transformations
$$
T_\theta^{\mathbf{B}}=
T_{\theta_n}^{\mathbf{B}_n}\circ\cdots\circ T_{\theta_1}^{\mathbf{B}_1}.
$$
For each variable $X_i$ in batch $\mathbf{B}_j$, the paper states
$$
X_i=T_{\theta_j}^{\mathbf{B}_j}(U_i\mid \mathbf{X}_{\mathbf{pa}_i}),
$$
so that each coordinate depends only on its own noise and its parents. This architecture is presented as causally consistent by construction, supports interventions and counterfactuals, and retains multilayer expressiveness unavailable to earlier causally consistent models [2412.12401].

The fairness application in that work makes the normalization interpretation explicit. On the German credit dataset, individual unfairness is measured by
$$
\frac{1}{n}\sum_{i=1}^{n}
\left|
Risk(X^{(i)}_{\mathrm{sex}=1})-Risk(X^{(i)}_{\mathrm{sex}=0})
\right|.
$$
Reported results are: SVM accuracy 73.00, F1 82.12, fairness 9.00; SVM$_{\mathrm{CF}}$ accuracy 72.60, fairness 4.10; and CCNF accuracy 75.80, F1 84.34, fairness 0.00. Here the causally normalized representation is the exogenous-noise space constrained not to encode forbidden causal paths from protected attributes to risk [2412.12401].

## 4. Entropy-budget normalization of causal influence in time series

In information-theoretic work on time-series causality, causal input normalization has a different meaning: it normalizes an information-flow rate by the full entropy balance of the recipient series. For a two-dimensional stochastic system, Liang defines the information flow from $X_2$ to $X_1$ as
$$
T_{2\to1}
=
-E\!\left[\frac{1}{\rho_1}\frac{\partial (F_1\rho_1)}{\partial x_2}\right]
+\frac{1}{2}E\!\left[\frac{1}{\rho_1}\frac{\partial^2 (g_{11}\rho_1)}{\partial x_2^2}\right],
$$
and decomposes the entropy rate of $X_1$ into self-dynamics, noise, and incoming information flow:
$$
\frac{dH_1}{dt}
=
\frac{dH_1^*}{dt}
+
\frac{dH_1^{\mathrm{noise}}}{dt}
+
T_{2\to1}.
$$
The normalized flow is then
$$
\tau_{2\to1}=\frac{T_{2\to1}}{Z_1},
\qquad
Z_1=T_{2\to1}+\frac{dH_1^*}{dt}+\frac{dH_1^{\mathrm{noise}}}{dt}.
$$
The point of the normalization is not symmetry between directions, but relative importance within the entropy budget of the recipient. The paper stresses that $Z_1\neq Z_2$ in general, so reverse flows should not share a common normalizer [1501.03548].

This formulation changes the interpretation of causality from an absolute rate to a causal share. In the bivariate AR(1) example with strong one-way coupling $(\alpha,\beta)=(0.5,0)$, the estimates are
$$
T_{2\to1}\approx 0.1481,\qquad T_{1\to2}\approx -0.0002,
$$
with normalized values
$$
\tau_{2\to1}\approx 17\%,\qquad \tau_{1\to2}\approx -0.03\%.
$$
A separate example yields nearly equal absolute flows, $|T_{2\to1}|\approx 0.13$ and $|T_{1\to2}|\approx 0.12$, but very different normalized magnitudes, $|\tau_{2\to1}|\approx 6.7\%$ and $|\tau_{1\to2}|\approx 13\%$. These examples are used to argue that identical absolute flows can have different relative importance in their respective entropy balances [1501.03548].

The multivariate generalization preserves the same intuition while modifying the normalizer. For a linear stochastic system,
$$
\frac{dH_i}{dt}
=
\frac{dH_i^*}{dt}
+\sum_{j\neq i}T_{j\to i}
+\frac{dH_i^{\mathrm{noise}}}{dt},
$$
and the paper defines
$$
Z_i=
\left|\frac{dH_i^*}{dt}\right|
+\sum_{j\neq i}|T_{j\to i}|
+\left|\frac{dH_i^{\mathrm{noise}}}{dt}\right|,
\qquad
\tau_{j\to i}=\frac{T_{j\to i}}{Z_i}.
$$
This yields a normalized vector of causal inputs for each node, including self-influence and noise. In the six-node VAR example, the reported normalized magnitudes include $|\tau_{6\to2}|=13.2\%$, $|\tau_{6\to5}|=12.5\%$, $\tau_{4\to5}=2.4\%$, and $\tau_{5\to4}=8.8\%$. The same framework is used to reconstruct directed graphs and self-loops under heavy noise and near synchronization [2104.11360].

## 5. Counterfactual normalization under dataset shift

Counterfactual Normalization addresses dataset shift by normalizing inputs with respect to stable and unstable causal paths. The setup assumes a causal DAG with observed variables $\mathbf{O}$, target $T$, optional selection variable $S$, and unobserved variables $\mathbf{U}$ representing domain-dependent confounding. Any active path to $T$ that passes through $S$ or a variable in $\mathbf{U}$ is treated as unstable. The procedure first removes observed variables with active unstable paths to $T$, producing a stable conditioning set $\mathbf{Z}$, and then attempts to retain some vulnerable variables by replacing them with counterfactual versions in which the effects of selected observed parents are removed [1808.03253].

The core operation is node-splitting. For a variable $Y$ and a subset of observed parents $\mathbf{P}\subseteq pa(Y)$, the method introduces a counterfactual node $Y(\mathbf{P}=\emptyset)$, reroutes the non-intervened parents into that node, and makes the counterfactual node a parent of the factual $Y$. Under additive SEMs, the counterfactual value is computed by subtracting the estimated contribution of the intervened parents. In the linear example
$$
Y=w_3T+w_4C+\varepsilon_Y,
$$
the normalized input is
$$
\hat Y(C=\emptyset)=Y-\hat w_4 C.
$$
The resulting feature retains the stable contribution of $T$ and noise while removing the unstable contribution of $C$ [1808.03253].

The simulated cross-hospital classification experiment makes the effect numerically explicit. Baseline logistic regression on $(Y,A,C)$ yielded source AUROC 0.95 and target AUROC 0.80. Counterfactual Normalization using $Y(A=\emptyset,C=\emptyset)$ yielded source AUROC 0.96 and target AUROC 0.97. The version that retained vulnerable variables, CFN (vuln), reported 0.97 on source and 0.92 on target. In the same experiment, the maximum Fisher’s ratio increased from 0.66 for the baseline to 3.13 for CFN, the intraclass/interclass distance ratio decreased from 0.10 to 0.02, and the MST boundary proportion decreased from 0.56 to 0.22 [1808.03253].

The sepsis application uses selection bias rather than only unobserved confounding. INR is modeled as
$$
Y_i=
\beta_0+\beta_1T_i+\bm{\beta}_2^\top\mathbf{A}_i
+\bm{\beta}_3^\top\mathbf{C}_i+\beta_4X_i
+\delta(T_i,\mathbf{A}_i,\mathbf{C}_i,X_i)+\varepsilon_i,
$$
with $\delta(\cdot)\sim\mathcal{GP}(0,\gamma^2K_{rbf})$ and $\varepsilon_i\sim\mathcal{N}(0,\sigma^2)$. The normalized feature is
$$
Y_i(\bm{\emptyset},\bm{\emptyset},\emptyset)
=
Y_i
-\hat{\bm{\beta}}_2^\top\mathbf{A}_i
-\hat{\bm{\beta}}_3^\top\mathbf{C}_i
-\hat\beta_4X_i.
$$
On biased training data, CFN was slightly worse than baseline, but on unbiased test data it improved AUPRC from 0.24 to 0.30 while baseline and CFN (vuln) remained at 0.24. This is the intended trade-off: reduced exploitation of unstable mechanisms in exchange for better transfer [1808.03253].

## 6. Online feature normalization, assumptions, and limits

A distinct contemporary usage appears in streaming monocular geometry, where Dynamic Feature Normalization is described as a lightweight, causal recurrent module that dynamically and robustly modulates feature statistics to maintain stable geometry over time. The paper’s empirical analysis traces temporal instability to fluctuations in latent feature statistics whose mean and variance directly determine predicted depth scale and shift. DyFN normalizes the current feature map,
$$
\mathcal{F}_t^{\mathrm{norm}}
=
\frac{\mathcal{F}_t-\boldsymbol{\mu}_{\mathcal{F}_t}}
{\boldsymbol{\sigma}_{\mathcal{F}_t}+\epsilon},
$$
updates a ConvGRU state
$$
\mathbf{h}_t=\mathrm{ConvGRU}(\mathcal{F}_t,\mathbf{h}_{t-1}),
$$
predicts stabilized statistics
$$
\hat{\boldsymbol{\sigma}}_t=\mathrm{Conv}_{1\times1}^{\sigma}(\mathbf{h}_t),
\qquad
\hat{\boldsymbol{\mu}}_t=\mathrm{Conv}_{1\times1}^{\mu}(\mathbf{h}_t),
$$
and reconstructs consistent features
$$
\mathcal{F}_t^{\mathrm{consistent}}
=
\hat{\boldsymbol{\sigma}}_t\cdot \mathcal{F}_t^{\mathrm{norm}}
+\hat{\boldsymbol{\mu}}_t.
$$
The normalization is causal in the online sense: only current and past frames are used [2605.25308].

This usage is not an SCM intervention method, but it preserves the central idea that normalization should respect causal access patterns. The module adds a mere 2\% additional parameters while keeping the backbone frozen. On video depth evaluation, the reported ScanNet results are MoGe v1 AbsRel 0.117 and $\delta<1.25=84.7$ versus DyFN AbsRel 0.073 and $\delta<1.25=96.6$; on KITTI, DyFN reports AbsRel 0.062 and $\delta_1=97.3$; on Bonn, AbsRel 0.044 and $\delta_1=98.4$. The paper also states that DyFN improves over prior streaming methods by up to 14\% and even outperforms heavier non-causal video baselines [2605.25308].

Across the broader literature, several assumptions recur. Flow-based causal normalization typically requires a known correct causal DAG or at least a known causal ordering; omitted confounders, wrong causal directions, or graph misspecification can invalidate counterfactual and interventional semantics. Counterfactual Normalization requires an additive SEM structure and observed parents whose effects can be subtracted. Time-series entropy-budget normalization is derived in closed form for linear stochastic systems with additive noise, and its estimators rely on covariance and derivative estimates. DyFN, by contrast, does not identify causal structure, but assumes that stabilizing the feature statistics that control scale and shift is sufficient to stabilize output geometry [2412.12401] [2306.05415] [1808.03253] [2104.11360] [2605.25308].

Taken together, these works show that causal input normalization is not a single algorithm but a methodological pattern. The normalized object may be a counterfactual input, an exogenous latent coordinate, a relative information-flow share, or a temporally stabilized feature map. What unifies them is the replacement of purely statistical normalization by a normalization rule anchored in causal structure, causal semantics, or causal access constraints [2106.06845] [1501.03548] [1808.03253] [2412.12401].

Source: https://www.emergentmind.com/topics/causal-input-normalization