Papers
Topics
Authors
Recent
Search
2000 character limit reached

NormReg: Perturbation-Based Normality Regularization

Updated 14 July 2026
  • Perturbation-Based Normality Regularization (NormReg) is a consistency mechanism that refines normal representations by ensuring perturbed and original embeddings remain close.
  • It complements teacher-guided score alignment by calibrating the student model against noisy teacher scores, thereby mitigating overfitting to limited labeled normal patterns.
  • The method uses stochastic feature masking on labeled nodes to tighten the normal cluster, yielding improved AUROC and AUPRC metrics in anomaly detection.

Searching arXiv for the cited papers to ground the article in the current literature. arXiv paper metadata:

  • (Zeng et al., 2 Oct 2025) — "Normality Calibration in Semi-supervised Graph Anomaly Detection" (published 2025-10-02)
  • (Triki et al., 2016) — "Stochastic Function Norm Regularization of Deep Networks" (published 2016-05-30)
  • (Chen et al., 2020) — "Stabilizing Differentiable Architecture Search via Perturbation-based Regularization" (published 2020-02-12) Perturbation-Based Normality Regularization (NormReg) is the representation-space regularization component of GraphNC, a teacher-student framework for semi-supervised graph anomaly detection (GAD) that was introduced to calibrate graph normality beyond the narrow patterns captured by labeled normal nodes alone (Zeng et al., 2 Oct 2025). Within GraphNC, NormReg does not replace teacher-guided score alignment; it complements it by enforcing perturbation consistency on labeled normal nodes, thereby encouraging more stable and compact normal representations under mild feature corruption. In operational terms, NormReg regularizes “graph normality” by requiring that a labeled normal node and its perturbed counterpart remain close in the student embedding space, which is intended to reduce the impact of inaccurate teacher scores and, in particular, reduce false positives caused by overfitting to limited labeled normal patterns (Zeng et al., 2 Oct 2025).

1. Position within semi-supervised graph anomaly detection

GraphNC is formulated for the semi-supervised GAD setting in which a subset of annotated normal nodes is available during training. The paper identifies a central limitation of this regime: the learned notion of normality is often confined to the labeled normal node set Vl\mathcal V_l, which inclines existing methods to overfit the observed normal patterns and to assign high anomaly scores to unlabeled-but-actually-normal nodes that differ from that subset (Zeng et al., 2 Oct 2025). The reported consequence is high detection error, especially high false positive rates.

NormReg is introduced specifically to address the failure mode that remains after teacher-guided score distillation. GraphNC first applies anomaly score distribution alignment (ScoreDA), which aligns the student’s anomaly scores with the score distribution produced by a pre-trained teacher model over all nodes. Because the teacher is often correct on most normal nodes and part of the anomaly nodes, this alignment tends to pull normal and abnormal scores toward opposite ends of the anomaly-score axis and increase score separability. However, the teacher is itself trained from limited labeled normal data and therefore inherits the same normality-characterization limitations. The paper is explicit that the teacher inevitably produces some inaccurate anomaly scores, and a student trained only by score alignment can fit those errors and become misled (Zeng et al., 2 Oct 2025).

NormReg is therefore framed as a compensatory mechanism. ScoreDA supplies teacher-guided supervision in anomaly-score space, whereas NormReg regularizes the student against overfitting to noisy teacher supervision by calibrating normality in representation space. The two modules are repeatedly described as complementary: ScoreDA uses the teacher’s useful information to shape the score distribution globally, while NormReg allows the student to self-refine its notion of normality through perturbation consistency on a trusted anchor set of labeled normal nodes.

2. Formal definition and objective

GraphNC uses a pre-trained teacher detector FT\mathcal F_{\mathcal T} with parameters Θ\Theta, producing anomaly scores

YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.

The student model FS\mathcal F_S has parameters Φ={Ω,ϕ}\Phi=\{\Omega,\phi\} and is implemented as a GNN encoder FGNN\mathcal F_{GNN} plus an MLP scorer FMLP\mathcal F_{MLP}. For node viv_i, the student representation and anomaly score are

hiS=FGNN(xi,G;Ω),yiS=FMLP(hiS;ϕ).{\bf h}_i^{\mathcal S} = \mathcal F_{GNN}({\bf x}_i,\mathcal G;\Omega), \qquad y_i^{\mathcal S} = \mathcal F_{MLP}({\bf h}_i^{\mathcal S};\phi).

The score-alignment loss is an MSE objective over all nodes:

FT\mathcal F_{\mathcal T}0

NormReg is defined over the labeled normal node set FT\mathcal F_{\mathcal T}1, where FT\mathcal F_{\mathcal T}2. The paper gives the NormReg loss as

FT\mathcal F_{\mathcal T}3

where FT\mathcal F_{\mathcal T}4 is the original student embedding and FT\mathcal F_{\mathcal T}5 is the student embedding of the perturbed view of the same labeled normal node (Zeng et al., 2 Oct 2025).

The surrounding text states that NormReg “minimizes the discrepancy” between the original and augmented representations. This suggests that the printed minus sign is a typographical inconsistency, because minimizing a negative squared distance would maximize discrepancy. A faithful conceptual reading is therefore a standard embedding-consistency penalty proportional to

FT\mathcal F_{\mathcal T}6

The complete GraphNC objective is

FT\mathcal F_{\mathcal T}7

where FT\mathcal F_{\mathcal T}8 controls the strength of NormReg relative to score alignment. At inference, the student score is used directly:

FT\mathcal F_{\mathcal T}9

with optimized student parameters Θ\Theta0 (Zeng et al., 2 Oct 2025).

The relevant symbols are specified as follows: Θ\Theta1 is an attributed graph; Θ\Theta2 is the node feature matrix; Θ\Theta3 is the adjacency matrix; Θ\Theta4 is the labeled normal node set; Θ\Theta5 is the unlabeled node set; Θ\Theta6 is the feature mask ratio; and Θ\Theta7 is squared Θ\Theta8 distance. The paper also states that there is no KL divergence, cosine similarity, expectation-based training loss, or adversarial perturbation term in the main NormReg objective.

3. Perturbation mechanism and consistency target

NormReg uses a node attribute-based masking mechanism rather than graph-structure corruption. The paper states that it “randomly masks a proportion Θ\Theta9 of the attributes on labeled normal nodes” to create an augmented graph view (Zeng et al., 2 Oct 2025). The perturbation therefore has four explicit properties: it is applied to node features or attributes, it is restricted to labeled normal nodes, it is stochastic, and it is implemented by feature masking.

The method description also specifies what the perturbation is not. NormReg does not define edge dropping, adjacency perturbation, message perturbation, Gaussian noise injection, parameter perturbation, or adversarial perturbation. Although the notation YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.0 appears in the algorithm, the textual description specifies only feature masking; a practical interpretation is therefore that the graph “augmentation” is induced by masked node attributes rather than by an independently defined structural corruption rule (Zeng et al., 2 Oct 2025).

Consistency is enforced on the student embeddings, not on anomaly scores, logits, or probabilities. The regularized quantity is the squared Euclidean discrepancy

YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.1

for nodes YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.2. The restriction to labeled normal nodes is a central design choice. The paper argues that only these nodes are known to be normal, so applying consistency to unlabeled nodes would risk compacting representations of potential anomalies and blurring the boundary between normal and abnormal. This is reinforced by an ablation variant, OT+ScoreDA+NormReg*, which applies NormReg to all nodes rather than only labeled normal nodes and performs worse than the default model (Zeng et al., 2 Oct 2025).

The perturbation-based view of “graph normality” is therefore local and invariance-based: if a labeled normal node undergoes a mild stochastic corruption of its attributes, its embedding should remain close to the original embedding. This defines a more stable normal region in latent space without directly imposing a score-level consistency constraint.

4. Training procedure and optimization

GraphNC follows a teacher-student pipeline that is effectively two-stage. First, a teacher model YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.3 is pre-trained using an existing semi-supervised GAD method. Second, the teacher is frozen and only the student is trained under the joint objective YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.4 (Zeng et al., 2 Oct 2025). The paper is explicit that there is no alternating teacher-student update, no momentum or EMA teacher, and no pseudo-labeling.

Algorithmically, the procedure tied to NormReg is:

  1. Obtain the teacher score distribution YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.5.
  2. Create an augmented feature matrix YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.6 by applying RandomMask to features of nodes in YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.7.
  3. Compute student representations on the original graph/features and on the perturbed graph/features.
  4. Compute student anomaly scores.
  5. Compute YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.8 and YT={y1T,y2T,…,yNT}.\mathcal Y^{\mathcal T}=\{y_1^{\mathcal T},y_2^{\mathcal T},\ldots,y_N^{\mathcal T}\}.9.
  6. Combine them into FS\mathcal F_S0 and update the student parameters FS\mathcal F_S1 by gradient descent.

The algorithm writes the original and augmented representations as

FS\mathcal F_S2

FS\mathcal F_S3

Optimization uses Adam. The default learning rates are FS\mathcal F_S4 for Photo and Reddit, and FS\mathcal F_S5 for Amazon, T-Finance, YelpChi, and Tolokers. The default regularization hyperparameters are FS\mathcal F_S6 and FS\mathcal F_S7, so 30% of attributes of labeled normal nodes are randomly masked by default (Zeng et al., 2 Oct 2025).

A plausible implication is that NormReg mainly shapes the encoder part of the student, because its loss acts directly on FS\mathcal F_S8 outputs, while ScoreDA additionally supervises the downstream scoring head. The paper’s wording supports this division of labor by stating that NormReg affects representation learning, whereas ScoreDA calibrates the anomaly scores.

5. Geometric and theoretical interpretation

NormReg is intended to make normal node representations more compact. The paper states that the augmentation simulates “diverse normal patterns that may be different from the ones derived directly from the labeled nodes,” and the consistency objective maps original and perturbed views of labeled normal nodes close together in latent space (Zeng et al., 2 Oct 2025). The geometric effect is described as a tightened normal cluster, reduced sensitivity of normal embeddings to feature perturbation, a smoother local normal manifold, and reduced intra-class variance among normal nodes.

The method does not explicitly repel anomalies. Its primary effect is to tighten the normal class rather than directly push anomalies away. Improved separation is therefore indirect: as the normal cluster becomes more compact, overlap between normal and abnormal score distributions is reduced. The paper describes this operationally as helping pull many normal-node anomaly scores closer together and toward the lower end of the anomaly-score distribution, thereby reducing false positives in particular (Zeng et al., 2 Oct 2025).

The theoretical discussion connects this representation-level regularization to score variance reduction. The paper states that minimizing NormReg together with ScoreDA leads to shrinking score variance in the normal class,

FS\mathcal F_S9

where Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}0 and Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}1 denote the teacher and student normal-class score variance, respectively.

For analysis, the normal embedding for node Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}2 under two views is modeled as

Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}3

where Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}4 is a latent normal prototype or center and Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}5 are perturbation noises. Under this model,

Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}6

Its expectation becomes

Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}7

under the paper’s independence and zero-mean assumptions. The stated conclusion is that minimizing the consistency loss reduces the variance of perturbation-induced deviations and compacts embeddings around Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}8 (Zeng et al., 2 Oct 2025).

The paper also characterizes this as reducing the average deviation of normal nodes to the normal prototype. The provided t-SNE visualizations are described as showing visibly tighter normal representation distributions when NormReg is used. This suggests that the method’s core geometry is cluster compactness and local invariance rather than contrastive separation by explicit negative pairs.

The principal empirical support for NormReg comes from GraphNC ablations. The comparison among OT, OT+ScoreDA, OT+NormReg, OT+NormReg-Finetune, OT+ScoreDA+NormReg*, and OT+ScoreDA+NormReg indicates that NormReg improves teacher-guided score alignment when used as the full GraphNC model (Zeng et al., 2 Oct 2025). In particular, comparing OT+ScoreDA with OT+ScoreDA+NormReg yields the following average gains:

Variant AUROC Avg AUPRC Avg
OT+ScoreDA 0.7274 0.3155
OT+ScoreDA+NormReg* 0.7356 0.3045
OT+ScoreDA+NormReg 0.7533 0.3610

Dataset-level improvements over ScoreDA alone are also reported for Amazon, T-Finance, Reddit, YelpChi, Tolokers, and Photo. The paper interprets these gains as evidence that ScoreDA improves the teacher but remains sensitive to inaccurate teacher scores, whereas NormReg mitigates that weakness (Zeng et al., 2 Oct 2025).

The worse performance of OT+ScoreDA+NormReg* is particularly consequential because it supports the design choice of applying consistency solely on labeled normal nodes. NormReg applied to all nodes is consistently inferior to the default model, which directly supports the claim that the trusted normal anchor set Φ={Ω,ϕ}\Phi=\{\Omega,\phi\}9 should define the latent normal manifold. The variants OT+NormReg and OT+NormReg-Finetune further show that NormReg alone can help somewhat on some datasets, but it is weaker and less stable than the joint use of ScoreDA and NormReg. The method is therefore presented as a complementary module rather than a standalone substitute for distillation.

Sensitivity analyses add two caveats. Performance is generally stable as FGNN\mathcal F_{GNN}0 varies, but too large an FGNN\mathcal F_{GNN}1 can slightly hurt on some datasets because overly strong consistency can lead to an over-compressed representation space and reduce discriminability. Likewise, different datasets prefer different masking strengths FGNN\mathcal F_{GNN}2; increasing FGNN\mathcal F_{GNN}3 helps some datasets but hurts others, suggesting that the useful perturbation magnitude depends on the variation present in normality (Zeng et al., 2 Oct 2025).

In relation to adjacent regularization traditions, the paper positions NormReg as a perturbation-based embedding consistency regularizer tailored to semi-supervised GAD with only normal labels. It does not benchmark directly against VAT, dropout consistency, edge perturbation consistency, or graph contrastive objectives. A broader function-space perspective is provided by “Stochastic Function Norm Regularization of Deep Networks” (Triki et al., 2016), which is conceptually relevant because it argues that parameter norms are not proper function norms and instead penalizes the weighted FGNN\mathcal F_{GNN}4 norm of the network output under a sampling distribution FGNN\mathcal F_{GNN}5. That method is not itself a perturbation-based local normality regularizer; it is best understood as a global, distribution-weighted output-energy penalty. A different partial analogue appears in “Stabilizing Differentiable Architecture Search via Perturbation-based Regularization” (Chen et al., 2020), which perturbs architecture parameters rather than inputs or embeddings and links perturbation robustness to Hessian-related smoothness in architecture space. Compared with these related lines, NormReg is distinguished by acting

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Perturbation-Based Normality Regularization (NormReg).