Papers
Topics
Authors
Recent
Search
2000 character limit reached

CCAD: Compressed Global Feature Anomaly Detection

Updated 1 January 2026
  • The paper introduces a novel two-stream architecture that combines global feature compression with latent diffusion reconstruction, achieving superior image and pixel-level AUCs.
  • The methodology integrates a fixed pretrained encoder and a U-Net-style diffusion model enhanced by cross-attention, enabling rapid convergence even under domain shifts.
  • Extensive experiments, including on the re-annotated DAGM 2007 dataset, demonstrate that CCAD delivers precise anomaly localization and robust performance across industrial benchmarks.

Compressed Global Feature Conditioned Anomaly Detection (CCAD) is an advanced paradigm unifying unsupervised representation-based and reconstruction-based approaches for industrial visual anomaly detection. CCAD exploits global dataset-level feature banks as explicit conditioning for a diffusion-based reconstruction model, facilitating robust feature extraction and efficient training, particularly under domain shift and limited anomalous supervision. The architecture incorporates a two-stage feature compression mechanism and integrates cross-attention-driven global conditioning into a U-Net-style latent diffusion network. Extensive experiments, including with a re-annotated DAGM 2007 dataset, establish superior convergence speed and state-of-the-art image/pixel-level AUCs across diverse benchmarks (Jin et al., 25 Dec 2025).

1. High-Level Framework

CCAD comprises two tightly coupled streams:

  • Global Feature Compression Stream: Extracts feature representations from all normal training images using a fixed pretrained encoder (e.g., ResNet), forming a large pool of dataset-level features. This pool undergoes compression, producing a compact, representative global feature bank.
  • Diffusion-Based Reconstruction Stream: Utilizes a latent diffusion model (U-Net backbone) to denoise a noisy latent vector back to a normal image distribution. The reconstruction is conditioned on local features from the input image and on the compressed global-feature bank.

During inference, anomalous images are reconstructed towards the learned normal distribution. Anomaly scores are derived by comparing features (e.g., cosine similarity) between the original and reconstructed images on both pixel and image levels.

2. Adaptive Global Feature Compression

Given normal training images X={xi}i=1N\mathcal{X} = \{x_i\}_{i=1}^N of size H×WH\times W, a fixed encoder F\mathcal{F} extracts dd-dimensional features from patches:

vn=F(xn),D={vn∈Rd}n=1M,M=N⋅⌊H/m⌋⋅⌊W/m⌋,v_n = \mathcal{F}(x_n),\quad \mathcal{D} = \{v_n \in \mathbb{R}^d\}_{n=1}^M, \quad M = N\cdot \lfloor H/m \rfloor \cdot \lfloor W/m \rfloor,

where mm is the patch downsampling factor. Since D\mathcal{D} can be excessively large, CCAD introduces a two-stage compression:

  • Coarse Feature Bank (CFB): A coreset-sampling operator S\mathcal{S} selects ξ\xi representative features (ξ≤1000\xi \leq 1000) from H×WH\times W0:

H×WH\times W1

  • Fine Feature Bank (FFB): At each training step, a batch-level subset H×WH\times W2 of size H×WH\times W3 is sampled. The trainable Fine Compression Module (FCM) H×WH\times W4 cross-attends between H×WH\times W5 (queries) and H×WH\times W6 (keys/values), yielding a refined bank H×WH\times W7 via

H×WH\times W8

where H×WH\times W9 are learned matrices. This multi-level compression efficiently condenses dataset-level priors for scalable and adaptive conditioning.

3. Diffusion Model Architecture and Conditioning

CCAD leverages a latent diffusion framework. An image F\mathcal{F}0 is mapped through an encoder F\mathcal{F}1 to latent F\mathcal{F}2. At diffusion timestep F\mathcal{F}3:

F\mathcal{F}4

The denoising network F\mathcal{F}5 predicts added noise, where F\mathcal{F}6 encodes conditional information.

Global-feature Conditioned Blocks (GCB) are inserted into the U-Net at multiple resolutions. Each GCB augments the ResBlock and self-attention with a cross-attention operation over either F\mathcal{F}7 or F\mathcal{F}8, depending on the variant. This design enables each U-Net stage to access holistic dataset statistics.

The three official variants are:

Variant Diffusion Space Global Conditioning
CCAD(V) Pixel-space Coarse bank F\mathcal{F}9, no fine
CCAD(C) Latent-space Coarse bank dd0
CCAD(F) Latent-space Trainable fine bank dd1

4. Training Objectives and Loss Functions

CCAD extends the standard denoising diffusion probabilistic model (DDPM) loss:

dd2

For CCAD(F), the condition dd3 includes local features and dd4:

dd5

For CCAD(C):

dd6

No additional regularizers are applied beyond the dd7 noise prediction term. Conditions can be augmented with ControlNet local features or global banks as required by the variant.

5. Reorganized and Re-Annotated DAGM 2007 Dataset

The original DAGM 2007 dataset supplied 15,000 normal and 2,100 defective synthetic texture images, labeled by coarse ellipses. CCAD authors re-annotated four challenging classes—defect, scratch, blur, and spots—with precise pixel-accurate masks. For each, 300 normal images were used for training, all defectives and 75 normals for testing.

Experiments confirm that the re-annotated masks yield almost identical image-level AUC, but improved pixel-level AUC for all methods, indicating that the finer masks correspond more exactly with true anomaly contours. Consequently, the refined DAGM supports more reliable benchmarking of pixel-wise anomaly localization (Jin et al., 25 Dec 2025).

6. Experimental Benchmarks and Quantitative Results

CCAD was evaluated on MVTec-AD, VisA, MVTec-3D, MVTec-Loco, MTD, and the new DAGM splits. Metrics include AUROC (image/pixel), maximal F1, and average precision.

Key quantitative results:

Method Image-AUC (MVTec-AD) Pixel-AUC (MVTec-AD) Epochs to Reference AUC
CCAD(V) 0.968 0.965 500–3000
DDAD 0.962 0.966 —
PatchCore 0.858 0.948 —
CCAD(C) 0.953–0.961 0.959–0.962 100
CCAD(F) 0.953–0.961 0.959–0.962 110
DiAD 0.950 0.954 200

Across VisA, MVTec-3D, and MVTec-Loco, CCAD consistently matches or outperforms prior SOTA. Ablation studies show that reducing coarse bank size to dd8 retains most AUC, indicating robust selection of informative global features by the network's cross-attention mechanism.

A plausible implication is that the inclusion of dataset-level statistics via compressed global banks not only provides stronger priors but also accelerates convergence, as seen in the reduction of epochs required to reach benchmark AUC values compared to existing methods.

7. Qualitative Analysis of Anomaly Localization

CCAD provides reconstructions that distinctly remove anomalies in the original input, resulting in heatmaps reflecting the cosine similarity between reconstructions and ground truth masks. Qualitative samples from MVTec-AD, VisA, MVTec-3D, and MVTec-Loco demonstrate that CCAD yields sharper anomaly localization than prior diffusion models such as DDAD and DiAD. In cases utilizing the re-annotated DAGM, generated anomaly maps align more accurately with the fine-grained ground-truth masks.

The ability to maintain precise and consistent localization further supports CCAD's advantage in industrial visual inspection and similar domains where both detection accuracy and fine localization are critical (Jin et al., 25 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Compressed Global Feature Conditioned Anomaly Detection (CCAD).