---
title: Causally-Guided Pairwise Transformer (CGPT)
url: https://www.emergentmind.com/topics/causally-guided-pairwise-transformer-cgpt
type: topic
---

# Causally-Guided Pairwise Transformer (CGPT)

The Causally-Guided Pairwise Transformer (CGPT) is a Transformer architecture for industrial multivariate time-series forecasting that integrates a known causal graph as an inductive bias and resolves the channel-dependent (CD) versus channel-independent (CI) trade-off through a pairwise modeling paradigm [2508.13111]. In this formulation, directed target-context pairs aligned with a causal graph are processed by shared, channel-agnostic modules whose parameter dimensions are independent of the number of variables, so that CGPT is CD at the pair level and CI-like across pairs [2508.13111]. The resulting architecture is designed for “any-variate” adaptability, multivariate-to-univariate forecasting, and digital twin settings in which explicit cross-variable interactions, lagged effects, and changing sensor configurations are all operationally relevant [2508.13111].

## 1. Conceptual basis and problem setting

CGPT is motivated by a stated tension in industrial multivariate time-series modeling. Channel-Dependent models learn a single function over all variables jointly and can capture explicit cross-variable interactions, multivariate lags, and non-linearities, but their parameterization typically scales with the number of channels \(C\), and any change in the sensor set or task forces architectural reconfiguration or retraining [2508.13111]. Channel-Independent models learn a shared function applied to each variable individually and therefore generalize well across datasets and are robust to noise and distribution shifts, but they are structurally blind to explicit cross-channel dependencies that matter in physically coupled systems [2508.13111].

CGPT addresses this by decomposing multidimensional data into directed target-context pairs \((i \to j)\) aligned with a causal graph \(G=(V,E)\), where each node corresponds to a channel and an edge \((i \to j)\in E\) indicates that channel \(i\) is a potential causal parent of channel \(j\) [2508.13111]. For a target \(j\), the parent set is defined as
\[
\mathrm{Pa}(j)=\{i \mid (i\to j)\in E\},
\]
and pair formation is restricted to these edges, which is central to both computational tractability and inductive bias [2508.13111].

The forecasting problem is posed on a multivariate time series \(X\in\mathbb{R}^{T\times C}\), optionally with exogenous covariates \(U\in\mathbb{R}^{T\times D_U}\) [2508.13111]. With context length \(L_{\mathrm{ctx}}\) and prediction horizon \(H_{\mathrm{pred}}\), the model maps a history to future targets as
\[
f_\theta:\mathbb{R}^{L_{\mathrm{ctx}}\times C}\to\mathbb{R}^{H_{\mathrm{pred}}\times |\mathcal{Y}|},
\]
where \(\mathcal{Y}\) is the set of predicted channels [2508.13111]. The reported use case emphasizes multivariate-to-univariate tasks because of computational feasibility and prevalence in industrial digital twin applications [2508.13111].

A central design claim is that the model enforces CD information flow at the pair level and CI-like generalization across pairs, yielding an architecture whose learned modules do not scale with \(C\) [2508.13111]. This is the technical sense in which CGPT is presented as an “any-variate” model.

## 2. Architecture and mathematical formulation

CGPT operates in three stages: channel-independent temporal encoding, pairwise interaction modeling, and target-wise aggregation and prediction [2508.13111].

In Stage 1, each channel \(v\) is processed independently using shared modules. RevIN normalization is applied per series to mitigate distribution shift, then patch tokenization segments \(x^{(v)}\in\mathbb{R}^{L_{\mathrm{ctx}}}\) into \(N_p\) non-overlapping patches of length \(P\) with stride \(S\), typically \(P=S=32\) [2508.13111]. A patch embedding layer produces token representations \(Z_0^{(v)}\in\mathbb{R}^{N_p\times d_{\mathrm{model}}}\), and a shared Transformer encoder with \(E_{\mathrm{layers}}=1\), heads \(=1\), \(d_{\mathrm{model}}=64\), and FFN \(=128\) yields encoded channel tokens \(Z_E^{(v)}\in\mathbb{R}^{N_p\times d_{\mathrm{model}}}\) [2508.13111]. The encoding step is written as
\[
Z_E^{(v)} = \mathrm{Encoder}(Z_0^{(v)}), \qquad v\in\{1,\dots,C\}.
\]

The self-attention inside this encoder uses standard scaled dot-product attention,
\[
\mathrm{Attention}(Q,K,V)=\mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V,
\]
with \(Q=ZW_Q\), \(K=ZW_K\), and \(V=ZW_V\) [2508.13111]. The paper notes explicitly that, as implemented, the causal graph \(G\) is not used to mask attention inside this encoder; causal guidance enters instead through pair selection in the next stage [2508.13111].

In Stage 2, the model constructs a pairwise influence representation for each directed edge \((i\to j)\in E\). First, encoder outputs are pooled:
\[
z^{(v)}=\mathrm{pool}(Z_E^{(v)})\in\mathbb{R}^{d_{\mathrm{model}}}.
\]
For each directed pair \((i\to j)\), CGPT computes
\[
p_{i\to j}=f_\theta(x^{(i)},x^{(j)}) := \mathrm{MLP}_{\mathrm{pair}}([z^{(i)};z^{(j)}]) \in \mathbb{R}^{d_{\mathrm{model}}},
\]
where the same pair module is reused for all pairs [2508.13111]. This shared pair module is the locus of channel-dependent interaction modeling: the target and one parent are jointly processed to produce an influence vector [2508.13111].

In Stage 3, influences from all parents of a target are aggregated additively. For target \(j\),
\[
r_j = b_j + \sum_{i\in \mathrm{Pa}(j)} p_{i\to j}, \qquad b_j := z^{(j)},
\]
and the prediction head maps this representation to the forecast,
\[
\hat{y}_j = h(r_j)\in\mathbb{R}^{H_{\mathrm{pred}}}.
\]
Training uses a regression objective, typically MSE, with MAE also reported:
\[
\mathcal{L}=\sum_{t\in\mathcal{T}}\sum_{j\in\mathcal{Y}}\sum_{h=1}^{H_{\mathrm{pred}}}\big(y_j(t+h)-\hat{y}_j(t+h)\big)^2.
\]
The decomposition into base representation and influences is learned implicitly; the paper states that no explicit regularizers on influences are added in the reported model [2508.13111].

This architecture has two notable properties. First, all learned modules are shared across channels and pairs, so the encoder, pair MLP, and prediction head remain channel-agnostic [2508.13111]. Second, causal structure is enforced not by masking token-token attention but by selecting which parent-target pairs are instantiated, a design distinction that separates CGPT from later graph-masked Transformer variants [2508.13111].

## 3. Channel-agnostic parameterization, algorithms, and complexity

The computational argument for CGPT follows directly from its pairwise decomposition. Let \(C\) be the number of channels, \(|E|\) the number of directed edges, \(K_j=|\mathrm{Pa}(j)|\), \(N_p\approx L_{\mathrm{ctx}}/P\) the number of patches, and \(d=d_{\mathrm{model}}\) [2508.13111]. The CI encoder over all channels has per-layer complexity \(O(C\cdot N_p\cdot d^2)\), the pairwise module over all edges has complexity \(O(|E|\cdot d^2)\), and aggregation plus prediction for target \(j\) costs \(O(K_j\cdot d + d\cdot H_{\mathrm{pred}})\) [2508.13111]. Naively modeling all pairs would scale as \(O(C^2)\), whereas graph sparsity reduces this to \(O(|E|)\) [2508.13111].

The model pipeline is reported as follows. For each window in the dataloader, RevIN is applied per channel, patch embeddings are computed, the shared encoder produces \(Z\), pooled channel vectors \(z\) are formed, pairwise influences are computed for each target and each parent in \(\mathrm{Pa}(j)\), the target representation \(r_j\) is aggregated, the head predicts \(\hat{y}_j\), and optimization proceeds with MSE under AdamW and cosine annealing [2508.13111]. The reported hyperparameters are patch size/stride \(32/32\), \(d_{\mathrm{model}}=64\), FFN \(=128\), one encoder layer, one attention head, batch size \(256\), \(100\) epochs, early stopping patience \(10\), and learning rate \(10^{-3}\) [2508.13111].

The term “any-variate adaptability” is tied to the fact that parameterization is independent of \(C\): adapting to a new dataset or task requires changing the channel set and the graph \(G\), not the architecture [2508.13111]. A plausible implication is that the model is intended to support operational environments in which sensor availability and target definitions change over time, provided the new configuration can still be expressed in terms of directed parent-target pairs.

The same paper also emphasizes that CGPT is computationally most natural for multivariate-to-univariate forecasting, since unrestricted multivariate target prediction would multiply the number of instantiated pairs and thereby raise runtime and memory costs [2508.13111]. In this sense, the architecture’s scalability depends not only on channel count but also on causal sparsity.

## 4. Empirical evaluation and reported behavior

CGPT is evaluated on two synthetic datasets and four industrial datasets: Synthetic Additive, Synthetic Interactive, ETTh1, Multi-Stage Factory, Amino Emissions, and Refial [2508.13111]. Two task settings are reported: long-term forecasting with context \(=96\) and horizon \(=96\), and one-step forecasting with context \(=96\) and horizon \(=1\) [2508.13111]. Metrics are MSE and MAE, reported as mean \(\pm\) std over five runs [2508.13111].

The baselines are DLinear as a CI model and an MLP baseline as a CD model, together with three CGPT variants: LeakyPairwise, StrictPairwise, and PureInfluence [2508.13111]. LeakyPairwise is the full CGPT, StrictPairwise restricts intra-module gradient flow by using a generic learnable embedding for the target’s influence path, and PureInfluence removes the target’s own history in the final prediction step [2508.13111].

For long-term forecasting, the reported highlights include the following [2508.13111]:

- On Synthetic Additive, LeakyPairwise achieves MAE/MSE \(=0.3165/0.1581\) with RevIN off, slightly better than DLinear \(0.3181/0.1598\) and comparable to MLP \(0.3177/0.1592\).
- On Synthetic Interactive, LeakyPairwise records \(0.8992/1.8751\), better than DLinear \(0.9129/1.8965\) and comparable to MLP \(0.9000/1.8807\).
- On ETTh1 with RevIN on, LeakyPairwise gives \(0.1843/0.0575\), while DLinear is slightly better on MSE at \(0.1782/0.0556\).
- On Multi-Stage Factory with RevIN on, LeakyPairwise yields \(0.0327/0.0022\), approximately matching DLinear \(0.0319/0.0021\).
- On Refial, with RevIN on, LeakyPairwise achieves the best MSE at \(0.0982\) versus MLP \(0.1047\) and DLinear \(0.1897\).

The paper’s interpretation is that, in long-horizon tasks, pairwise modeling of causal drivers provides a consistent advantage over CI baselines, especially in industrial datasets with lagged, non-linear dynamics, while remaining competitive with end-to-end CD baselines [2508.13111].

For one-step forecasting, the reported conclusion is different. On ETTh1 with RevIN on, DLinear dominates with MAE/MSE \(=0.0457/0.0039\), whereas LeakyPairwise records \(0.0472/0.0041\) [2508.13111]. On Synthetic Interactive, DLinear is best at \(0.3229/0.2131\), with LeakyPairwise around \(0.3415/0.2323\), and on Refial DLinear is approximately \(0.0038/\approx 0\), compared with LeakyPairwise at approximately \(0.0043/\approx 0\) [2508.13111]. The reported conclusion is that, for one-step predictions, strong autoregression dominates and CI models excel [2508.13111].

The ablation results are equally central to the definition of CGPT. Removing access to target history through PureInfluence causes severe degradation across datasets, which the paper presents as evidence that forecasting cannot be treated effectively as purely extrinsic regression [2508.13111]. By contrast, the difference between StrictPairwise and LeakyPairwise is marginal, suggesting that the primary performance driver is not fine-grained gradient routing inside the pair module but the presence of target history in the overall architecture [2508.13111]. The effect of causal guidance is discussed qualitatively: replacing the graph-guided pair set by fully connected pairs would inflate \(|E|\), runtime, and the risk of negative transfer [2508.13111].

## 5. Relation to causal guidance, pairwise masking, and adjacent Transformer families

Within the broader literature, CGPT occupies a specific point in the design space of causally informed Transformers. In CGPT as reported, causal guidance is used to select parent-target pairs, not to mask self-attention [2508.13111]. This differs from the DAG-aware Transformer for causal effect estimation, which encodes a user-specified DAG directly into the attention mechanism via a hard mask:
\[
M_{ij} =
\begin{cases}
0 & \text{if } M^{adj}_{ji}=1 \text{ or } i=j \\
1 & \text{otherwise}
\end{cases}
\]
and
\[
\mathbf{A}^{mask} = \mathbf{A} + \mathbf{M}\cdot(-\infty),
\]
thereby enforcing parent-only attention inside each head [2410.10044]. That model is aimed at ATE and CATE estimation rather than time-series forecasting, but it formalizes a “pairwise and causally guided” attention mechanism in a way that is conceptually adjacent to CGPT [2410.10044].

A second neighboring line comes from autoregressive time-series transformers used for causal discovery. “Transformer Is Inherently a Causal Learner” proves, under assumptions A1–A4 and regularity, that the gradient sensitivities of transformer outputs with respect to lagged inputs recover direct lagged parents through the score gradient energy
\[
H_{j,i}^{\ell}:=\mathbb{E}\!\left[\Big(\partial_{x_{j,t-\ell}}\log p(X_{i,t}\mid X_{<t})\Big)^2\right] > 0
\]
iff the edge \(j \stackrel{\ell}{\longrightarrow} i\) exists [2601.05647]. That paper culminates in a concrete blueprint for a Causally-Guided Pairwise Transformer in which an extracted adjacency \(M_{i,j,\delta}\in\{0,1\}\) is used to restrict attention and regularize the forecast objective [2601.05647]. This suggests a direct route from data-driven causal discovery to graph-guided pair construction, even though the reported CGPT paper assumes a known graph rather than learning one [2508.13111; 2601.05647].

A third related model is the Causally Guided Transformer (CGT) for multivariate time-series anomaly detection. CGT uses a time-lagged causal graph prior and a per-target hard parent mask \(\pi_i\in\{0,1\}^P\) built by PCMCI so that
\[
X^{(i)}_{t,c} = X_t \odot \pi_i,
\]
which zeroes non-parent inputs before attention and learning [2604.17998]. The paper states explicitly that CGT is “essentially the same idea” as a Causally-Guided Pairwise Transformer in spirit and mechanism, although it implements causal guidance at the feature level for each target variable’s forecasting block rather than by creating variable-lag pair tokens and masking the attention matrix across them [2604.17998]. It also provides an explicit attention-mask interpretation for a stricter pairwise variant:
\[
A=\mathrm{softmax}\!\left(\frac{QK^\top + M}{\sqrt{d_k}}\right),
\]
with \(M[u,v]=0\) for allowed causal parent interactions and \(M[u,v]=-\infty\) otherwise [2604.17998].

Taken together, these adjacent papers clarify a useful taxonomy. CGPT, in the narrow sense of [2508.13111], is a graph-guided pair-selection architecture with shared pair modules. DAG-aware Transformers and CGT instantiate harder forms of causal guidance through attention masks or feature masks [2410.10044; 2604.17998]. The causal-discovery paper supplies a theoretical basis for deriving such masks from trained autoregressive Transformers rather than assuming them a priori [2601.05647]. A plausible implication is that later “causally guided” Transformer variants can be read as architectural generalizations of the same underlying principle: causal structure restricts which pairwise interactions are permitted or emphasized.

## 6. Interpretability, digital twins, limitations, and future directions

CGPT is presented as a model for foundational digital twins because it operationalizes domain causal knowledge through sparse, directed influence paths while remaining adaptable to arbitrary sensor configurations without architectural changes [2508.13111]. The architecture decomposes prediction into a target base representation and additive pairwise influences, which the paper frames as interpretable components of system dynamics [2508.13111]. Influence vectors \(p_{i\to j}\) can be examined through their norms, their contributions to \(r_j\), and sensitivity analyses, while attention maps within the CI encoder can be inspected per channel to analyze temporal patterns [2508.13111]. In the Refial case, the paper specifically points to gas flow, door state, hull temperature, and exhaust temperature as channels whose influences can be examined for insights into control efficacy, lagged effects, and thermal inertia [2508.13111].

The principal limitation is reliance on a known causal graph \(G\) [2508.13111]. The paper states that incorrect or incomplete graphs may omit important drivers or include spurious ones, affecting performance [2508.13111]. It therefore proposes several future directions: pruning by regularization on influence magnitudes, data-driven graph refinement followed by expert validation, and learnable influence aggregation at Stage 3 to downweight spurious parents [2508.13111]. Additional proposed extensions include attention-based aggregation over parents,
\[
r_j = b_j + \sum_i \alpha_{i\to j} p_{i\to j},
\]
cross-attention at the pair level with graph-guided masks, self-supervised pretraining to encode causal proximity, probabilistic forecasting, and multi-hop pairs with counterfactual forecasting conditioned on planned control inputs [2508.13111].

The paper also notes computational trade-offs when \(|E|\) is large and recommends graph-guided sparsity, degree capping, edge sampling during training, and multi-hop selection with pruning [2508.13111]. These are practical consequences of the pairwise design rather than contradictions of it. In encyclopedia terms, CGPT should therefore be understood not as a generic Transformer with causal terminology added post hoc, but as a specific architecture whose defining features are graph-guided pair formation, shared channel-agnostic parameterization, additive influence aggregation, and a forecasting orientation tailored to industrial digital twin use cases [2508.13111].

In the broader development of causal Transformers, CGPT marks a distinct architectural answer to a recurring problem: how to preserve explicit cross-variable structure without binding the model irreversibly to a fixed channel set. The reported evidence indicates that this answer is most effective in long-term forecasting and industrial settings where sparse causal drivers matter, while one-step forecasting remains dominated by strong autoregression and favors simpler CI baselines [2508.13111].

Source: https://www.emergentmind.com/topics/causally-guided-pairwise-transformer-cgpt