Papers
Topics
Authors
Recent
Search
2000 character limit reached

Attention Adversarial Dual AutoEncoder

Updated 2 December 2025
  • The paper introduces ADAEN, a hybrid framework combining dual autoencoder architectures with an attention layer to prioritize salient features in anomaly detection.
  • It integrates adversarial training with ranking-based prioritization and active learning, using GAN-generated augmentations to reduce labeled data requirements.
  • Empirical results demonstrate significant improvements in anomaly scoring and ranking across diverse operating systems, highlighting its applicability in real-world security contexts.

The Attention Adversarial Dual AutoEncoder (ADAEN) is a hybrid neural anomaly detection framework introduced for highly imbalanced detection tasks such as advanced persistent threat (APT) identification in security provenance trace data. ADAEN integrates dual autoencoder architectures, an attention mechanism, adversarial learning, ranking-based prioritization, and a data-efficient active learning loop with generative augmentation to achieve superior anomaly prioritization with minimal labeled data requirements (Benabderrahmane et al., 25 Nov 2025).

1. Dual AutoEncoder Architecture

ADAEN consists of two feed-forward autoencoders (AEs), denoted AE₁ and AE₂. AE₁ functions as a generator focusing on capturing core data characteristics, while AE₂ serves as a complementary refiner/discriminator, promoting distinct but compatible latent representations. Formally, let X={x(i)}i=1m\mathcal{X} = \{x^{(i)}\}_{i=1}^m with x(i)∈Rdx^{(i)}\in\mathbb{R}^d. Each AE maps x∈Rdx\in\mathbb{R}^d to a latent code z∈Rkz\in\mathbb{R}^k and reconstructs it:

zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),

x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,2

Each network uses LeakyReLU activations, batch normalization, and dropout. The loss for each AE is its mean squared reconstruction error:

Lrecj=1∣X∣∑x∈X∥x−gϕj(fθj(x))∥22\mathcal{L}_{rec_j} = \frac{1}{|\mathcal{X}|} \sum_{x \in \mathcal{X}} \| x - g_{\phi_j}(f_{\theta_j}(x)) \|^2_2

Total reconstruction loss is a weighted combination:

Lrec=αLrec1+(1−α)Lrec2,α=0.5\mathcal{L}_{rec} = \alpha \mathcal{L}_{rec_1} + (1-\alpha)\mathcal{L}_{rec_2}, \quad \alpha=0.5

2. Attention Mechanism

ADAEN incorporates an attention layer between AE₁’s encoder and decoder to enhance focus on salient features or temporal slices of the latent code, z1∈RT×kz_1\in\mathbb{R}^{T \times k}. For each slice hi∈Rkh_i\in\mathbb{R}^k and context vector x(i)∈Rdx^{(i)}\in\mathbb{R}^d0, compute attention weights:

x(i)∈Rdx^{(i)}\in\mathbb{R}^d1

The attended summary x(i)∈Rdx^{(i)}\in\mathbb{R}^d2 is:

x(i)∈Rdx^{(i)}\in\mathbb{R}^d3

This vector x(i)∈Rdx^{(i)}\in\mathbb{R}^d4 is input to the decoder, enabling adaptive emphasis on features relevant for reconstruction, and consequently anomaly discrimination.

3. Adversarial Training Regime

ADAEN introduces an adversarial game where a discriminator x(i)∈Rdx^{(i)}\in\mathbb{R}^d5 (parameterized as a 3-layer MLP) aims to separate genuine data from reconstructed outputs of AE₁ and AE₂. The adversarial losses are:

Discriminator:

x(i)∈Rdx^{(i)}\in\mathbb{R}^d6

Generator (AE₁, AE₂):

x(i)∈Rdx^{(i)}\in\mathbb{R}^d7

The overall objective combines reconstruction and adversarial losses via hyperparameter x(i)∈Rdx^{(i)}\in\mathbb{R}^d8 (x(i)∈Rdx^{(i)}\in\mathbb{R}^d9):

x∈Rdx\in\mathbb{R}^d0

Optimization alternates between minimizing x∈Rdx\in\mathbb{R}^d1 (updating x∈Rdx\in\mathbb{R}^d2 with AE parameters fixed) and minimizing x∈Rdx\in\mathbb{R}^d3 (updating AE parameters with x∈Rdx\in\mathbb{R}^d4 fixed).

4. Anomaly Scoring and Ranking

During inference, ADAEN computes an anomaly score for each x∈Rdx\in\mathbb{R}^d5:

x∈Rdx\in\mathbb{R}^d6

Samples are sorted in descending order by x∈Rdx\in\mathbb{R}^d7, producing a ranked anomaly list. Evaluation prioritizes high-fidelity ranking, employing normalized discounted cumulative gain (nDCG):

x∈Rdx\in\mathbb{R}^d8

where x∈Rdx\in\mathbb{R}^d9 is the anomaly label and z∈Rkz\in\mathbb{R}^k0 is the ideal z∈Rkz\in\mathbb{R}^k1.

5. Active Learning with Data Augmentation

To address limited labeled anomaly data, ADAEN incorporates an active learning protocol:

  1. Train ADAEN on initial labeled set z∈Rkz\in\mathbb{R}^k2 (all normals).
  2. Score unlabeled pool z∈Rkz\in\mathbb{R}^k3 by z∈Rkz\in\mathbb{R}^k4.
  3. Set threshold z∈Rkz\in\mathbb{R}^k5 at the z∈Rkz\in\mathbb{R}^k6-th percentile of z∈Rkz\in\mathbb{R}^k7.
  4. Compute uncertainty z∈Rkz\in\mathbb{R}^k8 for z∈Rkz\in\mathbb{R}^k9.
  5. Select top-zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),0 points (minimum zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),1) and query an oracle for ground truth.
  6. For points confirmed as normal, augment using a GAN zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),2:
    • Train GAN on oracle-confirmed normals.
    • Generate synthetic normals zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),3.
  7. Update zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),4 new normals (real and synthetic).
  8. Retrain ADAEN and repeat steps up to zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),5 rounds.

Pseudocode (as in (Benabderrahmane et al., 25 Nov 2025)):

Lrecj=1∣X∣∑x∈X∥x−gϕj(fθj(x))∥22\mathcal{L}_{rec_j} = \frac{1}{|\mathcal{X}|} \sum_{x \in \mathcal{X}} \| x - g_{\phi_j}(f_{\theta_j}(x)) \|^2_25

6. Implementation and Hyperparameters

The ADAEN configuration is defined by several key hyperparameters and training details:

Symbol Meaning Default/Example Value
zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),6 Input dimension (dataset-dependent)
zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),7 Latent code size zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),8
zj=fθj(x)=σ(We(j)x+be(j)),z_j = f_{\theta_j}(x) = \sigma(W_e^{(j)}x + b_e^{(j)}),9 AE₁/AE₂ recon. weight x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,20
x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,21 Adversarial term weight x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,22
x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,23 Anomaly threshold percentile x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,24
x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,25 Oracle queries per round user-defined
x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,26 Max active learning iters 40
x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,27 Activation function LeakyReLU + batch-norm, dropout
x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,28 Discriminator net arch 3-layer MLP
x^j=gϕj(zj)=σ(Wd(j)zj+bd(j)),j=1,2\hat{x}_j = g_{\phi_j}(z_j) = \sigma(W_d^{(j)}z_j + b_d^{(j)}), \quad j=1,29 GAN generator/discriminator (see (Benabderrahmane et al., 25 Nov 2025))

Optimization uses Adam (lr=Lrecj=1∣X∣∑x∈X∥x−gϕj(fθj(x))∥22\mathcal{L}_{rec_j} = \frac{1}{|\mathcal{X}|} \sum_{x \in \mathcal{X}} \| x - g_{\phi_j}(f_{\theta_j}(x)) \|^2_20-Lrecj=1∣X∣∑x∈X∥x−gϕj(fθj(x))∥22\mathcal{L}_{rec_j} = \frac{1}{|\mathcal{X}|} \sum_{x \in \mathcal{X}} \| x - g_{\phi_j}(f_{\theta_j}(x)) \|^2_21, Lrecj=1∣X∣∑x∈X∥x−gϕj(fθj(x))∥22\mathcal{L}_{rec_j} = \frac{1}{|\mathcal{X}|} \sum_{x \in \mathcal{X}} \| x - g_{\phi_j}(f_{\theta_j}(x)) \|^2_22=0.5, Lrecj=1∣X∣∑x∈X∥x−gϕj(fθj(x))∥22\mathcal{L}_{rec_j} = \frac{1}{|\mathcal{X}|} \sum_{x \in \mathcal{X}} \| x - g_{\phi_j}(f_{\theta_j}(x)) \|^2_23=0.999), batch size 128, and early stopping with patience 10 on validation Lrecj=1∣X∣∑x∈X∥x−gϕj(fθj(x))∥22\mathcal{L}_{rec_j} = \frac{1}{|\mathcal{X}|} \sum_{x \in \mathcal{X}} \| x - g_{\phi_j}(f_{\theta_j}(x)) \|^2_24.

7. Application Context and Empirical Results

ADAEN targets extreme class-imbalance settings, exemplified by APT detection where attacks comprise as little as 0.004% of events in provenance trace databases (DARPA Transparent Computing). Empirical evaluation spans Android, Linux, BSD, and Windows datasets under two distinct attack scenarios. Adoption of the ranking- and active-learning-enhanced protocol yields significant improvements in detection rates and ranking metrics (nDCG), outperforming previous approaches in prioritizing true anomalies with reduced labeling overhead. A plausible implication is that such an architecture facilitates deployment in real-world security infrastructure with minimal manual annotation requirements (Benabderrahmane et al., 25 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Attention Adversarial Dual AutoEncoder.