Papers
Topics
Authors
Recent
Search
2000 character limit reached

ForensicsSAM: Robust IFDL Framework

Updated 3 July 2026
  • ForensicsSAM is a unified framework for image forgery detection and localization, integrating forgery and adversary experts to maintain high accuracy under adversarial conditions.
  • It employs LoRA-adapted modules within a frozen SAM backbone to inject forgery-specific knowledge and correct adversarial perturbations efficiently.
  • Experimental results demonstrate state-of-the-art performance with minimal F1 drop under targeted adversarial attacks across diverse datasets.

ForensicsSAM is a unified framework for robust, parameter-efficient image forgery detection and localization (IFDL) explicitly designed to resist adversarial attacks, built atop the Segment Anything Model (SAM) backbone. Utilizing both shared forgery experts and gated adversary experts, as well as an intrinsic adversary detector, ForensicsSAM achieves state-of-the-art (SOTA) performance on standard and adversarial benchmarks. The architecture enables simultaneous image-level forgery classification, pixel-level localization, and resilience to adversarial perturbations, maintaining significant robustness compared to existing PEFT-based approaches (Peng et al., 10 Aug 2025).

1. Architectural Overview

ForensicsSAM structurally extends the frozen SAM image encoder EθE_\theta by attaching two classes of adapter modules ("experts") and a lightweight adversary detector. The pipeline is as follows:

  • Input: X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}.
  • Adversary detector DaD_a, implemented as a ResNet-18, outputs score Sa=Da(X)S_a = D_a(X).
  • Gating variable Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau), with default threshold τ=0.5\tau=0.5.
  • SAM encoder EθE_\theta is augmented by:
    • Forgery experts ϕ\phi, always active, injected via LoRA updates into every Transformer block.
    • Adversary experts ψ\psi, conditionally activated when Ga=1G_a=1, also LoRA-injected in global attention and MLP layers.
  • The overall encoder is X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}0, followed by:
    • Forgery detector head X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}1: image-level output, X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}2.
    • Mask decoder X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}3: pixel-level localization, output mask X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}4.

The three-stage training protocol decouples: (i) forgery expert learning (on clean images), (ii) adversary detector learning (on clean+adversarial), and (iii) adversarial correction (adversary expert tuning on paired clean/adversarial inputs) (Peng et al., 10 Aug 2025).

2. Forgery Expert Injection Mechanism

Forgery experts are realized as low-rank "LoRA" adapters within each Transformer block's QKV projections and first MLP. At each layer, input tokens X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}5 (attention) and X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}6 (MLP) modulate the projections:

X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}7

where X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}8 are LoRA up/down matrices with rank X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}9. These adapters (DaD_a0) are trained to inject universal forgery-relevant artifact knowledge, enhancing both discrimination and localization. The optimization loss combines image-level binary cross-entropy (BCE) for forgery classification and pixel-level terms (Dice, BCE):

DaD_a1

This approach compensates for the lack of forgery-specific bias in the upstream (frozen) encoder, lifting both detection and localization metrics on clean benchmarks (Peng et al., 10 Aug 2025).

3. Adversarial Detection and Correction

A ResNet-18-based adversary detector DaD_a2 is trained to map input images to a scalar adversarial score:

DaD_a3

Supervision includes binary cross-entropy plus triplet loss to explicitly split feature distributions of clean and adversarial samples:

DaD_a4

During inference, if DaD_a5 (DaD_a6), adversary experts (DaD_a7) are activated. These are also LoRA adapters, injected at the same locations but gated. They operate according to:

DaD_a8

Adversarial correction aligns features from clean (DaD_a9) and adversarial (Sa=Da(X)S_a = D_a(X)0) passes at five intermediate global-transformer outputs:

Sa=Da(X)S_a = D_a(X)1

This mitigation strategy results in minimal F1 drop under adversarial input, effectively closing the clean-adv gap (Peng et al., 10 Aug 2025).

4. Adversarial Attacks and Robust Training

ForensicsSAM evaluates robustness to multiple adversarial attacks crafted exclusively via the upstream model, including MI-FGSM, PGN, BSR, and UMI-GRAT. Attacks maximize encoder representation deviation under Sa=Da(X)S_a = D_a(X)2 constraint:

Sa=Da(X)S_a = D_a(X)3

The training protocol proceeds:

  • Stage 1: (Sa=Da(X)S_a = D_a(X)4), trained on clean with Sa=Da(X)S_a = D_a(X)5.
  • Stage 2: Sa=Da(X)S_a = D_a(X)6, trained on clean and adversarial images with Sa=Da(X)S_a = D_a(X)7.
  • Stage 3: Sa=Da(X)S_a = D_a(X)8, updated by Sa=Da(X)S_a = D_a(X)9 with paired Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau)0, all other parameters frozen.

This framework ensures that both feature-alignment and explicit adversarial discrimination are optimized, with adversary experts dedicated only to adversarial samples due to gating.

5. Experimental Results

Empirical evaluation leverages a broad suite of IFDL benchmarks, training on CASIAv2, IMD20, FantasticReality, TamperedCR; and evaluating on 10 diverse test sets. Three principal metrics are reported: image-level detection accuracy (ACC), pixel-level localization F1 (permute-F1), and adversarial robustness (F1 drop under attack).

Summary (clean, average across 10 sets):

Model ACC F1
ForensicsSAM 0.817 0.753
Second best 0.759 0.659

Adversarial Robustness (Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau)1):

  • AutoSAM, FakeShield, SAFIRE: F1 drops Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau)260%.
  • ForensicsSAM: F1 drops only Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau)312%, retaining Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau)40.66.

Adversary Detection:

  • Detector accuracy (ACC): Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau)5–Ga=1(Sa>τ)G_a = \mathbf{1}(S_a>\tau)6 for all attacks/datasets.

Ablation experiments corroborate that without adversary experts, robustness collapses; with the full three-component design, ForensicsSAM remains performant for both clean and adversarial cases (Peng et al., 10 Aug 2025).

6. Comparative Significance and Practical Considerations

ForensicsSAM is the first PEFT-based SAM extension addressing all three facets of IFDL robustness: discrimination, localization, and adversarial resistance. Existing PEFT-SAMs—whether utilizing only downstream prompt-tuning or simple mask decoders—are vulnerable to upstream-based attacks; ForensicsSAM mitigates this via distributed expert-injection and a targeted detector module. The architecture remains parameter-efficient by localizing additional capacity within LoRA adapters.

Deployment requires only modest extra computation:

  • One forward pass through ResNet-18 for adversary detection;
  • Gated execution of adversary experts only when needed.

A plausible implication is that this approach could generalize to other vision PEFT frameworks facing similar upstream-attack vulnerabilities.

7. Code and Resources

The official implementation, pretrained weights, and reproducible scripts are at https://github.com/siriusPRX/ForensicsSAM (Peng et al., 10 Aug 2025). The repository provides inference and training instructions, with support for both clean and adversarial evaluation pipelines.

In summary, ForensicsSAM establishes a robust, unified, and extensible baseline for adversarially resilient image forgery detection and localization by combining multi-stage PEFT expert injection, effective adversary gating, and thorough empirical validation.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ForensicsSAM.