NormReg: Perturbation-Based Normality Regularization
- Perturbation-Based Normality Regularization (NormReg) is a consistency mechanism that refines normal representations by ensuring perturbed and original embeddings remain close.
- It complements teacher-guided score alignment by calibrating the student model against noisy teacher scores, thereby mitigating overfitting to limited labeled normal patterns.
- The method uses stochastic feature masking on labeled nodes to tighten the normal cluster, yielding improved AUROC and AUPRC metrics in anomaly detection.
Searching arXiv for the cited papers to ground the article in the current literature. arXiv paper metadata:
- (Zeng et al., 2 Oct 2025) — "Normality Calibration in Semi-supervised Graph Anomaly Detection" (published 2025-10-02)
- (Triki et al., 2016) — "Stochastic Function Norm Regularization of Deep Networks" (published 2016-05-30)
- (Chen et al., 2020) — "Stabilizing Differentiable Architecture Search via Perturbation-based Regularization" (published 2020-02-12) Perturbation-Based Normality Regularization (NormReg) is the representation-space regularization component of GraphNC, a teacher-student framework for semi-supervised graph anomaly detection (GAD) that was introduced to calibrate graph normality beyond the narrow patterns captured by labeled normal nodes alone (Zeng et al., 2 Oct 2025). Within GraphNC, NormReg does not replace teacher-guided score alignment; it complements it by enforcing perturbation consistency on labeled normal nodes, thereby encouraging more stable and compact normal representations under mild feature corruption. In operational terms, NormReg regularizes “graph normality” by requiring that a labeled normal node and its perturbed counterpart remain close in the student embedding space, which is intended to reduce the impact of inaccurate teacher scores and, in particular, reduce false positives caused by overfitting to limited labeled normal patterns (Zeng et al., 2 Oct 2025).
1. Position within semi-supervised graph anomaly detection
GraphNC is formulated for the semi-supervised GAD setting in which a subset of annotated normal nodes is available during training. The paper identifies a central limitation of this regime: the learned notion of normality is often confined to the labeled normal node set , which inclines existing methods to overfit the observed normal patterns and to assign high anomaly scores to unlabeled-but-actually-normal nodes that differ from that subset (Zeng et al., 2 Oct 2025). The reported consequence is high detection error, especially high false positive rates.
NormReg is introduced specifically to address the failure mode that remains after teacher-guided score distillation. GraphNC first applies anomaly score distribution alignment (ScoreDA), which aligns the student’s anomaly scores with the score distribution produced by a pre-trained teacher model over all nodes. Because the teacher is often correct on most normal nodes and part of the anomaly nodes, this alignment tends to pull normal and abnormal scores toward opposite ends of the anomaly-score axis and increase score separability. However, the teacher is itself trained from limited labeled normal data and therefore inherits the same normality-characterization limitations. The paper is explicit that the teacher inevitably produces some inaccurate anomaly scores, and a student trained only by score alignment can fit those errors and become misled (Zeng et al., 2 Oct 2025).
NormReg is therefore framed as a compensatory mechanism. ScoreDA supplies teacher-guided supervision in anomaly-score space, whereas NormReg regularizes the student against overfitting to noisy teacher supervision by calibrating normality in representation space. The two modules are repeatedly described as complementary: ScoreDA uses the teacher’s useful information to shape the score distribution globally, while NormReg allows the student to self-refine its notion of normality through perturbation consistency on a trusted anchor set of labeled normal nodes.
2. Formal definition and objective
GraphNC uses a pre-trained teacher detector with parameters , producing anomaly scores
The student model has parameters and is implemented as a GNN encoder plus an MLP scorer . For node , the student representation and anomaly score are
The score-alignment loss is an MSE objective over all nodes:
0
NormReg is defined over the labeled normal node set 1, where 2. The paper gives the NormReg loss as
3
where 4 is the original student embedding and 5 is the student embedding of the perturbed view of the same labeled normal node (Zeng et al., 2 Oct 2025).
The surrounding text states that NormReg “minimizes the discrepancy” between the original and augmented representations. This suggests that the printed minus sign is a typographical inconsistency, because minimizing a negative squared distance would maximize discrepancy. A faithful conceptual reading is therefore a standard embedding-consistency penalty proportional to
6
The complete GraphNC objective is
7
where 8 controls the strength of NormReg relative to score alignment. At inference, the student score is used directly:
9
with optimized student parameters 0 (Zeng et al., 2 Oct 2025).
The relevant symbols are specified as follows: 1 is an attributed graph; 2 is the node feature matrix; 3 is the adjacency matrix; 4 is the labeled normal node set; 5 is the unlabeled node set; 6 is the feature mask ratio; and 7 is squared 8 distance. The paper also states that there is no KL divergence, cosine similarity, expectation-based training loss, or adversarial perturbation term in the main NormReg objective.
3. Perturbation mechanism and consistency target
NormReg uses a node attribute-based masking mechanism rather than graph-structure corruption. The paper states that it “randomly masks a proportion 9 of the attributes on labeled normal nodes” to create an augmented graph view (Zeng et al., 2 Oct 2025). The perturbation therefore has four explicit properties: it is applied to node features or attributes, it is restricted to labeled normal nodes, it is stochastic, and it is implemented by feature masking.
The method description also specifies what the perturbation is not. NormReg does not define edge dropping, adjacency perturbation, message perturbation, Gaussian noise injection, parameter perturbation, or adversarial perturbation. Although the notation 0 appears in the algorithm, the textual description specifies only feature masking; a practical interpretation is therefore that the graph “augmentation” is induced by masked node attributes rather than by an independently defined structural corruption rule (Zeng et al., 2 Oct 2025).
Consistency is enforced on the student embeddings, not on anomaly scores, logits, or probabilities. The regularized quantity is the squared Euclidean discrepancy
1
for nodes 2. The restriction to labeled normal nodes is a central design choice. The paper argues that only these nodes are known to be normal, so applying consistency to unlabeled nodes would risk compacting representations of potential anomalies and blurring the boundary between normal and abnormal. This is reinforced by an ablation variant, OT+ScoreDA+NormReg*, which applies NormReg to all nodes rather than only labeled normal nodes and performs worse than the default model (Zeng et al., 2 Oct 2025).
The perturbation-based view of “graph normality” is therefore local and invariance-based: if a labeled normal node undergoes a mild stochastic corruption of its attributes, its embedding should remain close to the original embedding. This defines a more stable normal region in latent space without directly imposing a score-level consistency constraint.
4. Training procedure and optimization
GraphNC follows a teacher-student pipeline that is effectively two-stage. First, a teacher model 3 is pre-trained using an existing semi-supervised GAD method. Second, the teacher is frozen and only the student is trained under the joint objective 4 (Zeng et al., 2 Oct 2025). The paper is explicit that there is no alternating teacher-student update, no momentum or EMA teacher, and no pseudo-labeling.
Algorithmically, the procedure tied to NormReg is:
- Obtain the teacher score distribution 5.
- Create an augmented feature matrix 6 by applying
RandomMaskto features of nodes in 7. - Compute student representations on the original graph/features and on the perturbed graph/features.
- Compute student anomaly scores.
- Compute 8 and 9.
- Combine them into 0 and update the student parameters 1 by gradient descent.
The algorithm writes the original and augmented representations as
2
3
Optimization uses Adam. The default learning rates are 4 for Photo and Reddit, and 5 for Amazon, T-Finance, YelpChi, and Tolokers. The default regularization hyperparameters are 6 and 7, so 30% of attributes of labeled normal nodes are randomly masked by default (Zeng et al., 2 Oct 2025).
A plausible implication is that NormReg mainly shapes the encoder part of the student, because its loss acts directly on 8 outputs, while ScoreDA additionally supervises the downstream scoring head. The paper’s wording supports this division of labor by stating that NormReg affects representation learning, whereas ScoreDA calibrates the anomaly scores.
5. Geometric and theoretical interpretation
NormReg is intended to make normal node representations more compact. The paper states that the augmentation simulates “diverse normal patterns that may be different from the ones derived directly from the labeled nodes,” and the consistency objective maps original and perturbed views of labeled normal nodes close together in latent space (Zeng et al., 2 Oct 2025). The geometric effect is described as a tightened normal cluster, reduced sensitivity of normal embeddings to feature perturbation, a smoother local normal manifold, and reduced intra-class variance among normal nodes.
The method does not explicitly repel anomalies. Its primary effect is to tighten the normal class rather than directly push anomalies away. Improved separation is therefore indirect: as the normal cluster becomes more compact, overlap between normal and abnormal score distributions is reduced. The paper describes this operationally as helping pull many normal-node anomaly scores closer together and toward the lower end of the anomaly-score distribution, thereby reducing false positives in particular (Zeng et al., 2 Oct 2025).
The theoretical discussion connects this representation-level regularization to score variance reduction. The paper states that minimizing NormReg together with ScoreDA leads to shrinking score variance in the normal class,
9
where 0 and 1 denote the teacher and student normal-class score variance, respectively.
For analysis, the normal embedding for node 2 under two views is modeled as
3
where 4 is a latent normal prototype or center and 5 are perturbation noises. Under this model,
6
Its expectation becomes
7
under the paper’s independence and zero-mean assumptions. The stated conclusion is that minimizing the consistency loss reduces the variance of perturbation-induced deviations and compacts embeddings around 8 (Zeng et al., 2 Oct 2025).
The paper also characterizes this as reducing the average deviation of normal nodes to the normal prototype. The provided t-SNE visualizations are described as showing visibly tighter normal representation distributions when NormReg is used. This suggests that the method’s core geometry is cluster compactness and local invariance rather than contrastive separation by explicit negative pairs.
6. Empirical evidence, design choices, and related regularization perspectives
The principal empirical support for NormReg comes from GraphNC ablations. The comparison among OT, OT+ScoreDA, OT+NormReg, OT+NormReg-Finetune, OT+ScoreDA+NormReg*, and OT+ScoreDA+NormReg indicates that NormReg improves teacher-guided score alignment when used as the full GraphNC model (Zeng et al., 2 Oct 2025). In particular, comparing OT+ScoreDA with OT+ScoreDA+NormReg yields the following average gains:
| Variant | AUROC Avg | AUPRC Avg |
|---|---|---|
| OT+ScoreDA | 0.7274 | 0.3155 |
| OT+ScoreDA+NormReg* | 0.7356 | 0.3045 |
| OT+ScoreDA+NormReg | 0.7533 | 0.3610 |
Dataset-level improvements over ScoreDA alone are also reported for Amazon, T-Finance, Reddit, YelpChi, Tolokers, and Photo. The paper interprets these gains as evidence that ScoreDA improves the teacher but remains sensitive to inaccurate teacher scores, whereas NormReg mitigates that weakness (Zeng et al., 2 Oct 2025).
The worse performance of OT+ScoreDA+NormReg* is particularly consequential because it supports the design choice of applying consistency solely on labeled normal nodes. NormReg applied to all nodes is consistently inferior to the default model, which directly supports the claim that the trusted normal anchor set 9 should define the latent normal manifold. The variants OT+NormReg and OT+NormReg-Finetune further show that NormReg alone can help somewhat on some datasets, but it is weaker and less stable than the joint use of ScoreDA and NormReg. The method is therefore presented as a complementary module rather than a standalone substitute for distillation.
Sensitivity analyses add two caveats. Performance is generally stable as 0 varies, but too large an 1 can slightly hurt on some datasets because overly strong consistency can lead to an over-compressed representation space and reduce discriminability. Likewise, different datasets prefer different masking strengths 2; increasing 3 helps some datasets but hurts others, suggesting that the useful perturbation magnitude depends on the variation present in normality (Zeng et al., 2 Oct 2025).
In relation to adjacent regularization traditions, the paper positions NormReg as a perturbation-based embedding consistency regularizer tailored to semi-supervised GAD with only normal labels. It does not benchmark directly against VAT, dropout consistency, edge perturbation consistency, or graph contrastive objectives. A broader function-space perspective is provided by “Stochastic Function Norm Regularization of Deep Networks” (Triki et al., 2016), which is conceptually relevant because it argues that parameter norms are not proper function norms and instead penalizes the weighted 4 norm of the network output under a sampling distribution 5. That method is not itself a perturbation-based local normality regularizer; it is best understood as a global, distribution-weighted output-energy penalty. A different partial analogue appears in “Stabilizing Differentiable Architecture Search via Perturbation-based Regularization” (Chen et al., 2020), which perturbs architecture parameters rather than inputs or embeddings and links perturbation robustness to Hessian-related smoothness in architecture space. Compared with these related lines, NormReg is distinguished by acting