Papers
Topics
Authors
Recent
Search
2000 character limit reached

Patch-SVDD for Anomaly Detection

Updated 7 February 2026
  • The paper introduces Patch SVDD, an unsupervised framework that achieves 0.921 detection and 0.957 segmentation AUROC on the MVTec AD benchmark.
  • It employs dual-scale hierarchical encoding with 32×32 and 64×64 patches, using a pull-together loss and context-prediction to generate robust, fine-grained representations.
  • The approach fuses patch-level anomaly scores for both localization and segmentation, balancing local clustering against feature informativeness for superior performance.

Patch SVDD (Patch-level Support Vector Data Description) is an unsupervised anomaly detection and segmentation framework for images, designed to enable fine-grained localization of anomalies by extending the deep learning variant of SVDD to operate at the patch level and incorporating a self-supervised learning component. The method achieves state-of-the-art results on the MVTec AD benchmark, substantially improving both anomaly detection and segmentation performance relative to prior approaches (Yi et al., 2020).

1. Mathematical Foundation

Patch SVDD evolves from classical SVDD, which seeks the minimal-volume hypersphere that contains the embeddings of normal samples in feature space. In the deep SVDD model, standard kernel mappings are replaced with neural encoders, typically yielding the objective:

LSVDD=∑i=1N∥fθ(xi)−c∥2L_{\text{SVDD}} = \sum_{i=1}^N \| f_\theta(x_i) - c \|^2

where fθ(x)f_\theta(x) is the learned embedding, and cc is the centroid of normal embeddings.

A naïve patch-based extension directly applies this principle to overlapping image patches; however, heterogeneity among patches (e.g., object, background, texture) leads to high intra-class variation, making a single-centroid approach unsuitable for dense anomaly localization.

Patch SVDD circumvents this by introducing a “pull-together” loss between spatially adjacent patches:

LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_2

where p′p' is a neighbor of pp in the 3×3 grid, thereby encouraging locally similar patch embeddings to cluster without enforcing a global unimodal structure.

To ensure meaningful, non-collapsed representations, Patch SVDD integrates a self-supervised context-prediction task. Given two patches p,p2p, p_2 from the 3×3 neighborhood with ground-truth relative position y∈{1,...,8}y \in \{1, ..., 8\}, a classifier CϕC_\phi predicts yy from their embedding difference:

fθ(x)f_\theta(x)0

The total loss for training is:

fθ(x)f_\theta(x)1

2. Architecture and Self-Supervision

The encoder fθ(x)f_\theta(x)2 consists entirely of convolutional layers (no biases), followed by LeakyReLU activations with slope fθ(x)f_\theta(x)3. Patch SVDD employs a two-level hierarchical encoder, summarized as follows:

  • "Small" encoder (fθ(x)f_\theta(x)4): receptive field 32, operates on fθ(x)f_\theta(x)5 patches.
  • "Big" encoder (fθ(x)f_\theta(x)6): processes fθ(x)f_\theta(x)7 patches by subdividing into four fθ(x)f_\theta(x)8 sub-patches, embedding each with fθ(x)f_\theta(x)9, and aggregating via concatenation and cc0 convolutions.
  • Embedding dimensionality cc1.

The self-supervised classifier cc2 is a two-layer MLP with 128 hidden units, operating on embedding differences. To suppress spurious cues, color channels are randomly perturbed during patch sampling.

3. Training Methodology

Training is conducted exclusively on normal images resized to cc3:

  • For cc4, extract overlapping cc5 patches (stride 16).
  • For cc6, extract cc7 patches (stride 4).
  • Each iteration samples a patch cc8, a neighbor cc9 (for LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_20), and another local patch LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_21 (for LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_22).
  • Optimization uses Adam (LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_23), learning rate LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_24, batch size 256, and 50 epochs.
  • No geometric augmentations (e.g., flips, rotations) are used.
  • The hyperparameter LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_25 is selected per class; smaller values (LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_26) for object classes, larger values (LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_27) for textures.

4. Anomaly Detection and Segmentation Pipeline

After training, all normal patches from the training set are encoded and stored for nearest-neighbor searches, forming databases LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_28 and LSVDD′=∥fθ(p)−fθ(p′)∥2L_{\text{SVDD}}' = \| f_\theta(p) - f_\theta(p') \|_29.

For a test image:

  • Overlapping patches are extracted with the same strides and embedded via both encoders.
  • Patch anomaly score is computed as p′p'0, independently at small and big scales.
  • Pixel-level anomaly maps p′p'1 and p′p'2 are derived by distributing patch scores to pixels and averaging over coverage. They are fused via element-wise multiplication: p′p'3.
  • The global image anomaly score is p′p'4.

This design supports both image-level and fine-grained segmentation ROC analyses.

5. Experimental Results

Patch SVDD was evaluated on the MVTec AD dataset (15 industrial classes, comprising both objects and textures). Metrics are per-class AUROC for both detection and segmentation.

Method Detection AUROC Segmentation AUROC
Deep SVDD (ICML ’18) 0.592 –
GEOM (NeurIPS ’18) 0.672 –
GANomaly (ACCV ’18) 0.762 –
ITAE (arXiv ’19) 0.839 –
Patch SVDD (ours) 0.921 0.957
L₂-AE – 0.804
SSIM-AE – 0.818
VE-VAE (CVPR ’20) – 0.861
VAE Proj (ICLR ’20) – 0.893

Patch SVDD achieves +9.8% detection and +7.0% segmentation AUROC improvement over the prior best entries.

6. Functional Analysis and Ablation Insights

  • Replacing the single-center loss with "pull-together" (p′p'5) improves AUROC, and further addition of the context-prediction (p′p'6) yields the highest performance.
  • Ablation studies indicate that object classes, characterized by high intra-class patch variation, benefit more from self-supervised context prediction, while texture classes are less sensitive to this term.
  • Feature visualizations (t-SNE) show single-modality clusters without p′p'7, and semantic, multi-modal clusters when both losses are used. The lowest intrinsic feature dimension occurs when using both components.
  • Hierarchical multi-scale encoding (combining p′p'8 and p′p'9 scales) outperforms either scale alone and single-level multi-branch architectures, suggesting that shared sub-encoders induce regularization and an inductive bias.
  • The selection of pp0 balances local clustering against feature informativeness; optimal values differ per class.
  • Embedding dimensionality exhibits diminishing returns beyond pp1.
  • Surprisingly, nearest-neighbor anomaly detection on random CNN features or even raw pixel patches can suffice for certain classes, as for one-layer convolutions, Euclidean distance in feature and pixel space are closely related.

7. Limitations and Potential Extensions

Patch SVDD requires maintaining large databases of normal patch embeddings and performing approximate nearest-neighbor search at inference time, which can be both memory- and compute-intensive. Hyperparameter tuning of pp2 is manual and dataset-dependent. Extending the approach to additional scales, or incorporating further self-supervised objectives (e.g., rotation or jigsaw prediction), are plausible routes for increased robustness and generalization (Yi et al., 2020).

Code and pretrained models are publicly available, facilitating adaptation and reproduction for diverse anomaly detection pipelines.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Patch-SVDD for Anomaly Detection.