---
title: 'FLaG: A Multi-Domain Overview'
url: https://www.emergentmind.com/topics/flag
type: topic
---

# FLaG: A Multi-Domain Overview

Searching arXiv for papers on “FLaG/FLAG” to ground the article and verify the relevant usages across fields.
FLaG, or FLAG, appears in recent arXiv literature in several unrelated senses. In network coding it refers to **flag codes**, that is, sets of sequences of nested subspaces of a vector space over the finite field $\mathbb{F}_q$ [2207.01997]. In machine learning and statistics it names **Fast Label-Adaptive Aggregation** for federated multi-label classification [2302.13571], **Frequency-Domain Latent Attention Gating** for token aggregation [2606.08191], **Foundation model representation with Latent diffusion Alignment via Graph** for spatial gene expression prediction [2605.18055], and **Flexible and Accurate Methods for Estimation and Inference of Gaussian Graphical Models** [2306.17584]. The shared label therefore denotes a family of field-specific constructions rather than a single unified method.

## 1. Terminology and disciplinary scope

The principal arXiv uses of FLaG/FLAG can be organized as follows.

| Usage | Research area | Defining idea |
|---|---|---|
| FLaG [2207.01997], [2311.06060] | Network coding | Flag codes, distance vectors, equivalence, and automorphisms |
| FLAG [2302.13571] | Federated learning | Label-adaptive aggregation for multi-label classification |
| FLaG [2606.08191] | Representation learning | Frequency-domain latent attention gating for token aggregation |
| FLAG [2605.18055] | Spatial transcriptomics | Diffusion-based structured prediction with graph conditioning and GFM alignment |
| FLAG [2306.17584] | Statistical network modeling | Precision-matrix estimation and inference for Gaussian graphical models |

Outside these named constructions, **flag** also retains its standard mathematical and physical meanings. In tropical geometry, the **flag Dressian** is introduced as a tropical analogue of the partial flag variety and is linked to valuated flag matroids, flags of projective tropical linear spaces, and coherent flag matroidal subdivisions [2005.13727]. In fluid-structure interaction, a **flag near a free surface** denotes a flexible plate in axial flow, and varying the Froude number yields rigidly-confined flutter, a resonance regime, and softly-confined flutter [2005.08667].

## 2. FLaG as flag codes in network coding

In the network-coding literature, a **flag** in $\mathbb{F}_q^n$ is a strictly nested sequence of subspaces
$$
\{0\}\subsetneq \mathcal{F}_1 \subsetneq \mathcal{F}_2 \subsetneq \cdots \subsetneq \mathcal{F}_r \subsetneq \mathbb{F}_q^n,
$$
and its type is the dimension sequence $(\dim(\mathcal{F}_1),\dots,\dim(\mathcal{F}_r))$. A **flag code** of type $(t_1,\dots,t_r)$ is a nonempty set
$$
C \subseteq \mathcal{F}_q((t_1,\dots,t_r),n).
$$
The special case emphasized is the **full flag variety** $\mathcal{F}_q(n)$, where the type is $(1,2,\dots,n-1)$, so a full flag contains one subspace in every intermediate dimension [2207.01997].

The metric structure is built from the **injection distance**
$$
d_I(U,V)=\max\{\dim U,\dim V\}-\dim(U\cap V),
$$
which for equal-dimensional subspaces of dimension $k$ becomes
$$
d_I(U,V)=k-\dim(U\cap V).
$$
For two flags
$$
F=(F_1,\dots,F_r), \qquad F'=(F'_1,\dots,F'_r),
$$
the **flag distance** is
$$
d_f(F,F')=\sum_{i=1}^r d_I(F_i,F'_i),
$$
and the finer invariant is the **distance vector**
$$
d(F,F')=\big(d_I(F_1,F'_1),\dots,d_I(F_r,F'_r)\big).
$$
For full flags on $\mathbb{F}_q^n$, the distance vector has length $n-1$, and the maximum possible full-flag distance is
$$
D^n=\left\lfloor \frac{n^2}{4}\right\rfloor,
$$
which is sharp.

The central structural theorem characterizes exactly which vectors occur. A vector
$$
(\delta_1,\dots,\delta_{n-1})
$$
is a valid distance vector for some pair of full flags if and only if, after setting $\delta_0=\delta_n=0$, it satisfies
$$
\delta_i\in\{\delta_{i-1}-1,\ \delta_{i-1},\ \delta_{i-1}+1\}
\qquad \text{for all } i=1,\dots,n.
$$
Thus the sequence changes by at most $1$ at each step, starts at $0$, and ends at $0$. This is exactly the combinatorics of a **Motzkin path** of length $n$. The bijection is defined by reading increments:
$$
p_i=
\begin{cases}
U & \text{if } \delta_i-\delta_{i-1}=1,\\
H & \text{if } \delta_i-\delta_{i-1}=0,\\
D & \text{if } \delta_i-\delta_{i-1}=-1,
\end{cases}
$$
with $U=(1,1)$, $H=(1,0)$, and $D=(1,-1)$.

This yields the main enumerative statement: **the number of possible distance vectors for the full flag variety $\mathcal{F}_q(n)$ is the $n$-th Motzkin number $M_n$**. The first Motzkin numbers are
$$
1,1,2,4,9,21,51,127,\dots
$$
The same bijection refines the metric by identifying the flag distance with the **area under the Motzkin path**:
$$
A(\Psi(v))=\delta_1+\cdots+\delta_{n-1}.
$$
Hence vectors with fixed total distance $d$ correspond exactly to Motzkin paths of length $n$ and area $d$, counted by
$$
T(n,d)=\#\{p\in\mathcal{M}_n : A(p)=d\}.
$$
For a full flag code with minimum distance $d_f(C)$, the number of possible distance vectors at that minimum distance is $T(n,d_f(C))$. In the maximum-distance case $d_f(C)=D^n$, there is only one possible distance vector, because there is only one Motzkin path of maximal area. The paper also isolates the **disjoint** case, where no component subspaces coincide and the associated vectors have no zero entries; these correspond to **elevated Motzkin paths**. When two flags never share consecutive subspaces, the paths have no horizontal steps on the $x$-axis and are counted by the **Riordan numbers** [2207.01997].

## 3. Equivalence of flag codes, projected codes, and broader flag geometry

A later development studies when two flag codes should be regarded as the same up to linear or semilinear change of coordinates. For a flag code
$$
\mathcal{C}\subseteq \mathcal{F}_q((t_1,\dots,t_r),n),
$$
the $i$-th **projected code** is the constant-dimension code
$$
\mathcal{C}_i=\{F_i \mid (F_1,\dots,F_r)\in\mathcal{C}\}\subseteq \mathcal{G}_q(t_i,n).
$$
These projected codes collect the prescribed-dimensional subspaces appearing in each slot of the flags. If $A\in \mathrm{GL}_n(\mathbb{F}_q)$, or more generally $(A,\sigma)\in \Gamma\mathrm{L}_n(\mathbb{F}_q)$, the action is applied componentwise to flags, leading to the definitions of **linear equivalence**, **semilinear equivalence**, and the automorphism groups $\operatorname{Aut}(\mathcal{C})$ and $\operatorname{SAut}(\mathcal{C})$ [2311.06060].

Equivalence of flag codes always implies equivalence of all projected codes under the same group element. The converse is false in general, because the way the projections are assembled into nested tuples matters. The paper therefore identifies classes for which projected-code data suffice. A central notion is that of a **subspace-inclusion-closed** (**SIC**) flag code. For a generating family $\mathcal{C}_1\times\cdots\times\mathcal{C}_r$, the set
$$
\mathcal{C}= (\mathcal{C}_1\times\cdots\times\mathcal{C}_r)\cap \mathcal{F}_q((t_1,\dots,t_r),n)
$$
is the unique SIC flag code generated by that family, and every other generated flag code is contained in it. The paper then defines a code to be **determined by its projected codes** if it is the only flag code generated by $\mathcal{C}_1\times\cdots\times\mathcal{C}_r$. The characterization is multiplicity-theoretic:
$$
\mathcal{C}\text{ is determined by its projected codes} \iff
\mathcal{C}\text{ is SIC and every flag has a subspace of multiplicity }1.
$$

This criterion yields the main reduction theorem. If $\mathcal{C}$ is determined by its projected codes and another flag code $\mathcal{C}'$ satisfies
$$
\mathcal{C}'_i=\mathcal{C}_i\cdot A \quad\text{for all }i,
$$
or the semilinear analogue, then
$$
\mathcal{C}'=\mathcal{C}\cdot A.
$$
Within this class, equivalence of projected constant-dimension codes is enough to recover equivalence of the flag codes themselves. The automorphism theory also simplifies: in general
$$
\operatorname{Aut}(\mathcal{C}) \subseteq \bigcap_{i=1}^r \operatorname{Aut}(\mathcal{C}_i),
\qquad
\operatorname{SAut}(\mathcal{C}) \subseteq \bigcap_{i=1}^r \operatorname{SAut}(\mathcal{C}_i),
$$
but for SIC flag codes these inclusions are equalities.

The paper introduces two concrete families. A flag code is **increasing** if every subspace in $\mathcal{C}_i$ contains a unique subspace from $\mathcal{C}_{i-1}$; it is **decreasing** if each subspace in $\mathcal{C}_{i-1}$ is contained in a unique subspace in $\mathcal{C}_i$. Increasing or decreasing codes are SIC, are determined by their projected codes, and therefore admit equivalence reduction to the projected codes. For certain maximum-distance and optimum-distance flag codes, explicit type-vector conditions guarantee this behavior [2311.06060].

The broader mathematical setting of flags extends beyond coding. In tropical geometry, the **flag Dressian**
$$
FlDr(r_1,\dots,r_k;n)
$$
is defined as the tropical prevariety of the tropicalized equations cutting out the flag variety. The central theorem gives a four-way correspondence between points on the flag Dressian, valuated flag matroids, coherent flag matroidal subdivisions, and flags of projective tropical linear spaces. The same work proves that all valuated flag matroids on ground set up to size $5$ are realizable, and gives a $6$-element example where realizability fails [2005.13727].

## 4. FLAG as Fast Label-Adaptive Aggregation in federated learning

In federated learning, **FLAG** stands for **Fast Label-Adaptive Aggregation** and is designed specifically for **multi-label classification** rather than the much more common multi-class setting. The paper argues that standard aggregation methods such as **FedAvg** are not sufficient when the task is multi-label, because they ignore how labels co-occur, how frequent each label is on each client, and how unevenly label distributions vary across clients. The framework therefore addresses two gaps simultaneously: an **experimental gap**, because prior client simulations often borrow Dirichlet-based or random splitting from multi-class learning, and an **algorithmic gap**, because standard aggregation does not use client label composition [2302.13571].

The client simulation component is **Clustering-based Multi-label Data Allocation** (**CMDA**). Each sample is represented by its binary label vector $y_i\in\{0,1\}^L$, and the paper applies **k-modes clustering**, using the binary label vectors as clustering features and setting the number of clusters equal to the number of simulated clients $N$. The reported procedure is: represent each sample’s label vector in binary form, use these label vectors as clustering features, set the number of clusters equal to the number of simulated clients, run k-modes on the training labels, use the learned cluster centers to assign both training and validation data to clients, and build each client’s dataset from its assigned cluster. This is intended to produce **label-distribution skew** and data-size skew in a way that is more realistic for multi-label data.

The aggregation component computes a label-aware client weight:
$$
label\_weight = \left\{ \sum_{l=1}^{N_l}\left(\sum_{i=1}^{N_i^c} y_i^c\right)^\alpha,\ 1 \leq c \leq N_c \right\}.
$$
Here $N_c$ is the number of clients, $N_l$ the number of labels, $N_i^c$ the number of samples on client $c$, $y_i^c$ the binary label vector of the $i$-th sample on client $c$, and $\alpha$ a hyperparameter controlling the tradeoff between **label distribution** and **label occurrence**. The interpretation given is: $\alpha=0$ makes the weight depend only on label distribution wideness, while larger $\alpha$ makes label occurrence more important; experiments use $0\le \alpha\le 1$. The intended mechanism is a weighted model fusion rule where the weights come from local label statistics.

The workflow is explicit. CMDA first splits the centralized multi-label dataset into client datasets. Each client trains locally for several epochs, computes local label statistics, sends them to the server without exposing raw labels or data, uploads model parameters, and the server aggregates model updates using FLAG. The updated global model is then sent back to clients, and training repeats for multiple communication rounds.

The evaluation uses **MS-COCO 2014** as a multi-label image-classification benchmark, partitioned into **10 clients**. The backbone is **TResNet** with **Asymmetric Loss (ASL)**, **Adam**, **OneCycleLR**, batch size **128**, learning rate **$10^{-4}$**, weight decay **$10^{-4}$**, **4 epochs** of local training per communication round, and **40 epochs** total training. Baselines include centralized **TRresNet (G)**, local **TRresNet (L)**, **FedAvg**, **Per-FedAvg**, **pFedHN**, **Personal BN**, **Personal classifier**, **PFADET**, **KT-pFL**, and **FLAG-Aug**, which combines FLAG with **FedMix**. The metrics are **AmAP**, **WmAP**, and **GmAP**, together with training-epoch efficiency, communication-round efficiency, and convergence speed to a target performance. The target threshold is defined as **80\% of centralized performance**, corresponding to **48\% mAP**.

The main reported results are:
- **FedAvg**: AmAP = **47.9\%**, WmAP = **31.6\%**, GmAP = **50.3\%**
- **FLAG**: AmAP = **50.2\%**, WmAP = **31.2\%**, GmAP = **54.5\%**
- **FLAG-Aug**: AmAP = **50.9\%**, WmAP = **29.2\%**, GmAP = **54.8\%**

Relative to FedAvg, FLAG improves **AmAP by 4.8\%** and **GmAP by 8.3\%**; FLAG-Aug improves **AmAP by 6.3\%** and **GmAP by 8.9\%**. The paper further claims that FLAG needs **less than 50\% of the training epochs and communication rounds** required by state-of-the-art methods to reach comparable or better mAP, and can be **up to 2× faster** than other FedAvg-based aggregation methods. Results are stable for $\alpha=0.1$ to $0.8$, with best performance at **$\alpha=0.3$**, which is used as default. The paper also notes limitations: CMDA has randomness in clustering-based simulation, the aggregation weighting is static rather than dynamic, and experiments are limited to MS-COCO [2302.13571].

## 5. FLaG as Frequency-Domain Latent Attention Gating

In representation learning, **FLaG** stands for **Frequency-Domain Latent Attention Gating** and is presented as a **plug-in token aggregation module** designed to replace or augment standard pooling methods when a model must compress a variable-length token sequence into a sample-level representation. The argument is that token aggregation itself is often a bottleneck: standard pooling methods such as mean, max, last-token, and attention pooling operate only in the original token domain, whereas FLaG re-expresses the sequence in spectral coordinates, summarizes the spectrum with latent queries, reweights channels, reconstructs enhanced time-domain tokens, and only then applies final pooling [2606.08191].

For input
$$
X \in \mathbb{R}^{T \times D},
$$
with optional binary mask $m\in\{0,1\}^T$, FLaG applies the **real FFT** along the sequence dimension:
$$
X = \mathrm{rFFT}(X \odot m) \in \mathbb{C}^{F \times D},
\qquad
F = \left\lfloor \frac{T}{2} \right\rfloor + 1.
$$
The complex spectrum is represented as
$$
F = [\Re(X) \mid \Im(X)] \in \mathbb{R}^{F \times 2D},
$$
which preserves both magnitude and phase information. It then introduces $L$ learnable latent queries
$$
Q \in \mathbb{R}^{L \times 2D},
$$
and computes
$$
A = \mathrm{MHA}(Q, F, F) \in \mathbb{R}^{L \times 2D},
$$
followed by residual, normalization, and feed-forward processing:
$$
H = \mathrm{LayerNorm}(Q + \mathrm{Dropout}(A)),
$$
$$
S = \mathrm{LayerNorm}\big(H + \mathrm{Dropout}(\mathrm{FFN}(H))\big).
$$
Averaging across latent slots gives
$$
s = \frac{1}{L}\sum_{l=1}^{L} S_l \in \mathbb{R}^{2D},
$$
and the gate is
$$
g = \sigma(\mathrm{MLP}(s)) \in [0,1]^{2D}.
$$
The residual channel-wise modulation is
$$
F' = F \odot (1 + g),
$$
after which the model reconstructs the complex spectrum and applies the inverse FFT:
$$
X' = \mathrm{irFFT}\big(F^{(R)} + iF^{(I)}\big) \in \mathbb{R}^{T \times D}.
$$
Finally,
$$
z = \mathrm{Pool}(X', m),
$$
and the default choice is **max pooling**.

The method is evaluated across three domains: antimicrobial peptide activity prediction with **ESM2**, image classification with **ResNet18** on **CIFAR-10** and **CIFAR-100**, and text classification with **RoBERTa** on **IMDB** and **GLUE**. The clearest gains are reported for AMP prediction with the smaller **ESM2-8M** backbone. On **E. coli**, FLaG achieves RMSE **0.562** and Recall@50 **18.6**; on **S. aureus**, RMSE **0.545** and Recall@50 **18.0**, with the latter close to best max pooling at **18.4**. Compared to mean pooling, the RMSE improves from **0.578 → 0.562** on **E. coli** and from **0.577 → 0.545** on **S. aureus**. With **ESM2-35M**, results are more mixed, but FLaG achieves the best Recall@50 on **S. aureus** at **18.6**.

On vision benchmarks, FLaG reaches **96.01\%** on **CIFAR-10** and **77.20\%** on **CIFAR-100**. On CIFAR-100, the best baseline is mean pooling at **76.78\%**, so the reported gains are **+0.42 points** over mean pooling and **+0.55 points** over attention pooling; FLaG also has the **lowest variance** there. On **IMDB**, FLaG achieves **94.08\%**, tying mean pooling at **94.08\%**. On **GLUE**, mean pooling is best on MNLI, attention pooling is best on SST-2, and FLaG is characterized as competitive rather than dominant.

The diagnostic analyses on AMP tasks support a consistent picture. **Band knockouts** show that **low-frequency band $b=0$ contributes the most overall**, with strongest perturbation effects in middle-to-late transformer layers and higher-frequency contributions being more sample-specific. **Gate summaries** show that the gate **broadly amplifies the spectrum**, with post-gate energy higher across bands and low-frequency bands remaining most energetic before and after gating. **Residue perturbations** show that mean pooling is comparatively flat, whereas max pooling, attention pooling, and FLaG preserve more differentiated residue-response profiles. **Latent-query readouts** show sample-specific attention patterns with mild query-wise differences. **Structure-proxy stratification** shows that higher-helix peptides exhibit **stronger average band sensitivity** across both bacteria. Ablations indicate that **FFT + gate** captures a large part of the gain, while the full FLaG block is often best overall. The paper notes several limitations, including the assumption of regularly sampled, padded token sequences, possible padding-dependent spectral artifacts, additional computational overhead, and more mixed results on some text tasks [2606.08191].

## 6. FLAG as Foundation model representation with Latent diffusion Alignment via Graph

In spatial transcriptomics, **FLAG** denotes **Foundation model representation with Latent diffusion Alignment via Graph**. The task is to predict, from a whole-slide H\&E image with spatially sampled spots, each spot’s high-dimensional gene expression vector. For each spot $s$, the paper defines a tissue coordinate $u_s\in\mathbb{R}^2$, a visual feature $v_s\in\mathbb{R}^{d_v}$, and a gene-expression vector $x_s\in\mathbb{R}^G$. The set of spots forms a graph $G=(V,E)$ whose edges encode local microenvironmental relations based on physical proximity, histology similarity, and molecular similarity. The paper argues that most prior methods treat each gene at each spot as an independent scalar regression problem and therefore optimize pointwise criteria such as **MSE** or **PCC** while failing to preserve **gene-gene correlation structure** and **gene-spatial autocorrelation** [2605.18055].

The conceptual move is to treat the task as **structured distribution modeling** rather than deterministic pointwise regression. Diffusion is used to model the conditional distribution of expressions, but the paper identifies a major obstacle called the **Gene Dimension Curse**. In joint node-edge diffusion, both node states $X$ and edge states $A$ are denoised simultaneously; as gene dimension $G$ grows, the optimization rapidly deteriorates. The paper describes this through correlation concentration and gives the lower-bound statement
$$
E_{\text{joint}(G)} - E_{\text{node}} \ge \Omega(G).
$$
A plausible implication is that explicit generation of graph edges together with high-dimensional gene outputs becomes increasingly unstable as the target gene panel grows.

FLAG addresses this by decoupling spatial structure from gene generation. Instead of using the graph as a generative target, it uses a **Spatial Graph Encoder** to summarize topology, a **gene diffusion backbone** to generate expressions conditioned on that spatial context, and **Gene Foundation Model alignment** to preserve gene-gene fidelity. The factorization is written as
$$
\hat{X} = \epsilon(X_t \mid H_{\text{spatial}}, t),
$$
where $H_{\text{spatial}}$ is a compact spatial context embedding. The fixed tissue graph is constructed from
$$
W_{\text{dist}(i,j)}=\exp\!\left(-\frac{\|u_i-u_j\|^2}{\sigma^2}\right),
\qquad
W_{\text{img}(i,j)}=\mathrm{CosSim}(v_i,v_j),
$$
and the edge attribute is
$$
C_{e,ij} = [W_{\text{dist}(i,j)},\, W_{\text{img}(i,j)}].
$$
A **Graph Transformer** then computes
$$
H_{\text{spatial}} = \mathrm{GraphEncoder}(C_v, C_e).
$$

The graph encoder uses **AdaLayerNorm** and edge-modulated attention. Time embedding and pooled conditions are fused as
$$
z = \mathrm{MLPfuse}([\mathrm{temb}, \mathrm{Pool}(C_v), \mathrm{Pool}(C_e)]),
$$
with adaptive normalization
$$
H_x' = \mathrm{AdaLN}(H_x,z) = (1+\gamma_x(z)) \odot \mathrm{LN}(H_x) + \beta_x(z),
$$
and the static FLAG attention logits are
$$
S_{ij} = (q_i k_j^\top)\odot \big(1+\alpha \cdot \mathrm{Linear}(C_{e,ij})\big) + \gamma \cdot \mathrm{Linear}(C_{e,ij}).
$$
The spatial output is projected to the gene diffusion conditioner and passed to a **latent diffusion transformer (DiT)**. The denoising model is
$$
\epsilon_0(X_t)=\mathrm{DiT}(X_t\mid C_{\text{graph}}, t).
$$
The diffusion backbone uses the standard VE-SDE form
$$
x_t = x_0 + \sigma(t)z, \qquad z\sim \mathcal{N}(0,I),
$$
with score-matching objective
$$
L_{\text{score}} = \mathbb{E}\left\| s_\theta(x_t,t,c) - \nabla_{x_t}\log p_t(x_t\mid x_0) \right\|^2.
$$

To preserve gene semantics, the method aligns intermediate DiT features to frozen embeddings from pretrained **Gene Foundation Models**, including **Geneformer**, **scGPT**, and **CellPLM**, with **Geneformer** as the default/strongest setting. If $F\in\mathbb{R}^{G\times d_e}$ is the fixed per-gene embedding matrix and $H^{(k)}$ is an intermediate DiT representation, the alignment loss is
$$
L_{\text{align}} = 1 - \frac{\langle \mathrm{MLP}(H^{(k)}), F\rangle}{\|\mathrm{MLP}(H^{(k)})\|_2 \|F\|_2 + \epsilon},
$$
and the full objective is
$$
L_{\text{total}} = L_{\text{score}} + \lambda_{\text{align}} L_{\text{align}}.
$$

Because pointwise metrics are insufficient, the paper introduces two structural measures. **Gene Structural Correlation** (**GSC**) evaluates preservation of gene-gene regulatory structure. After standardizing $X$ and forming the gene-gene correlation matrix
$$
C = \frac{1}{N-1}\tilde{X}^\top \tilde{X},
$$
GSC is the correlation between the upper-triangular entries of the ground-truth and predicted matrices:
$$
\mathrm{GSC} = \mathrm{Corr}(V_{\mathrm{GT}}, V_{\mathrm{Pred}}).
$$
**Spatial Structural Correlation** (**SSC**) evaluates preservation of spatial autocorrelation using Moran’s $I$. With a symmetric $k$-nearest-neighbor graph,
$$
W_{ij}=1 \quad \text{if } j\in N_k(i) \text{ or } i\in N_k(j), \text{ else } 0,
$$
and $k=8$, Moran’s statistic for gene $g$ is
$$
I(g)=\frac{N}{S_0}\frac{z_g^\top W z_g}{z_g^\top z_g},
\qquad
S_0=\sum_{i,j}W_{ij},
$$
and
$$
\mathrm{SSC} = \mathrm{Corr}(I_{\mathrm{GT}}, I_{\mathrm{Pred}}).
$$

Experiments are conducted on **HEST-1k** cohorts—**HER2ST**, **KIDNEY**, and **PRAD**—with a **7:2:1** slide-level split, against baselines including **HisToGene**, **BLEEP**, **TRIPLEX**, **Stem**, and **STFlow**. FLAG is reported to be competitive on PCC/MSE while improving structural fidelity. On **HER2ST**, it achieves **PCC 0.6835**, **MSE 0.7342**, **GSC 0.8926**, and **SSC 0.6386**; on **KIDNEY**, **PCC 0.3917**, **MSE 1.2112**, **GSC 0.8713**, and **SSC 0.3409**; on **PRAD**, **PCC 0.5853**, **MSE 1.3771**, **GSC 0.8775**, and **SSC 0.7510**. The Gene Dimension Curse experiments show that **Joint Node-Edge Diffusion** performs well for small $G$ but collapses sharply as $G$ increases, whereas FLAG remains stable. Ablations further show that removing diffusion, GFM alignment, or the spatial graph degrades structural fidelity, and alignment to an intermediate GFM layer works best, with strongest GSC around **$l=8$**. Downstream analyses indicate better recovery of pathway co-expression structure, best overlap with ground-truth DEGs, and improved spatial domain identification. The paper notes limitations, including iterative diffusion inference, the current 2D formulation, and open questions on zero-shot cross-tissue generalization [2605.18055].

## 7. FLAG as Flexible and Accurate Methods for Estimation and Inference of Gaussian Graphical Models

In statistics, **FLAG** stands for **Flexible and Accurate Methods for Estimation and Inference of Gaussian Graphical Models**. The setting is a $p$-dimensional Gaussian vector
$$
z \sim N(0,\Sigma), \qquad \Theta=\Sigma^{-1},
$$
where the precision matrix encodes conditional dependence:
$$
z_i \perp z_j \mid z_{-(ij)} \iff \theta_{ij}=0.
$$
The corresponding graph $G=(V,E)$ has vertices $V=\{z_1,\dots,z_p\}$ and edges
$$
E=\{(z_i,z_j)\mid \Theta_{ij}\neq 0\}.
$$
The paper positions FLAG against penalized-likelihood, conditional-regression, and Bayesian methods, emphasizing accurate estimation without sparsity assumptions on $\Theta$, element-wise inference, computational efficiency, and extensions to multiple graphs and covariate adjustment [2306.17584].

The core construction is **pairwise conditional regression**. For each pair $(i,j)$, let
$$
Y=[Z_{\cdot i},Z_{\cdot j}] \in \mathbb{R}^{n\times 2},
\qquad
X=[Z_{\cdot i},Z_{\cdot j}]^c \in \mathbb{R}^{n\times(p-2)}.
$$
Conditional Gaussianity gives
$$
Y \mid X \sim N(X\beta,\Theta_{aa}^{-1}),
$$
where $a=\{i,j\}$ and $\Theta_{aa}$ is the corresponding $2\times 2$ precision block. Rather than estimating $\beta$ by sparse regression, FLAG uses a **random effects model**
$$
Y = X\beta + \epsilon,
$$
with
$$
\beta_{k\cdot}^T \sim N(0,\Gamma_\beta), \qquad \epsilon_{k\cdot}^T \sim N(0,\Gamma_\epsilon),
$$
where
$$
\Gamma_\beta=
\begin{bmatrix}
\sigma_1^2 & \tau\\
\tau & \sigma_2^2
\end{bmatrix},
\qquad
\Gamma_\epsilon=
\begin{bmatrix}
\sigma_3^2 & \eta\\
\eta & \sigma_4^2
\end{bmatrix}.
$$
The pairwise precision block is then recovered from
$$
\widehat{\Theta}_{aa}=\widehat{\Gamma}_\epsilon^{-1}.
$$
After integrating out $\beta$, the model becomes
$$
\operatorname{vec}(Y)\sim N(0,\Omega^{-1}),
\qquad
\Omega=\Gamma_\beta\otimes XX^T+\Gamma_\epsilon\otimes I_n,
$$
and estimation proceeds by maximizing the incomplete-data log-likelihood.

Graph recovery is formulated as hypothesis testing on the residual covariance. For each pair, the null hypothesis is
$$
H_0:\eta=0.
$$
Using asymptotic normality of the MLE,
$$
\sqrt{n}(\widehat{\Gamma}-\Gamma^*) \xrightarrow{d} N(0,I^{-1}),
$$
the paper constructs a **Wald test**
$$
W=\frac{(\eta-\eta_0)^2}{\operatorname{var}(\eta)}, \qquad \eta_0=0,
$$
and also a **likelihood ratio test**
$$
-2\big[\mathcal{L}(\Gamma_0)-\mathcal{L}(\Gamma)\big].
$$
After computing p-values for all pairs, graph recovery uses Benjamini–Hochberg FDR control or Bonferroni correction for FWER.

Two computational engines are developed. The first is a **parameter-expanded EM (PX-EM)** algorithm based on the expanded model
$$
Y=\delta X\beta+\epsilon.
$$
The E-step computes the Gaussian posterior of $\beta$, while the M-step updates $\Gamma_\beta$, $\Gamma_\epsilon$, and $\delta$, then rescales $\Gamma_\beta$ by $\delta^2$ and resets $\delta$ to $1$. The second is a **minorize-maximization (MM)** algorithm accelerated through eigendecomposition and low-rank updates. Since all $\binom{p}{2}$ pairwise regressions must be solved, the paper exploits
$$
XX^T = ZZ^T - Z_{\{ij\}}Z_{\{ij\}}^T,
$$
a rank-$2$ correction, together with the Woodbury identity and the matrix determinant lemma, to update inverses and determinants efficiently.

Two extensions are emphasized. **FLAG-Meta** performs joint estimation across multiple graphs at the **partial-correlation level**,
$$
\rho^{(k)}_{ij} = -\frac{\theta^{(k)}_{ij}}{\sqrt{\theta^{(k)}_{ii}\theta^{(k)}_{jj}}}
= \frac{\eta^{(k)}}{\sigma_3^{(k)}\sigma_4^{(k)}},
$$
testing whether groups share the same partial correlation and combining estimates by inverse-variance weighting when appropriate. **FLAG-CA** incorporates covariates directly through
$$
Y=\Upsilon \zeta + X\beta + \epsilon,
$$
so that fixed effects and random effects are estimated jointly rather than in a two-stage residualization pipeline.

The reported empirical conclusions are that FLAG is robust to scaling, provides accurate estimation without sparsity assumptions, achieves strong graph recovery and FDR control, and supports multi-graph borrowing of strength. Applications include gene expression in the human brain, term association in university websites, and stock prices in the U.S. financial market. In the stock-price analysis, rolling-window estimates capture the spike in network instability during the Covid-19 market crash and the return to stability afterward more consistently than competing methods [2306.17584].

Source: https://www.emergentmind.com/topics/flag