Papers
Topics
Authors
Recent
Search
2000 character limit reached

SD2AIL: Synthetic Demonstrations for AIL

Updated 28 December 2025
  • SD2AIL is an adversarial imitation learning framework that employs diffusion models to synthesize expert-like trajectories for robust reward inference and policy optimization.
  • The framework integrates pseudo-expert generation with prioritized expert demonstration replay, effectively augmenting scarce expert datasets and improving sample efficiency.
  • Empirical results on MuJoCo tasks demonstrate that SD2AIL outperforms baselines, achieving higher stability and performance even in challenging low-data regimes.

SD2AIL (“Synthetic Demonstrations to Adversarial Imitation Learning”) is an adversarial imitation learning (AIL) framework that leverages diffusion models to generate synthetic, expert-like demonstrations for reward inference and policy optimization. SD2AIL addresses the challenge of limited expert trajectory data by augmenting small expert datasets with high-quality synthetic samples (pseudo-experts) generated via a conditional denoising diffusion probabilistic model, thus improving AIL performance and stability even in low-data regimes. This methodology is integrated into a discriminator’s learning process and further facilitated by a prioritized expert demonstration replay (PEDR) strategy, enabling scalable and robust imitation learning from sparse demonstrations (Li et al., 21 Dec 2025).

1. Background and Motivation

AIL achieves policy learning by training a discriminator, DD, to distinguish between expert and agent-generated (policy) trajectories, while the generator policy πθ\pi_\theta seeks to fool the discriminator, as in Generative Adversarial Imitation Learning (GAIL). AIL methods typically require many high-quality expert trajectories for reliable reward inference and stable agent training. However, expert data acquisition is often costly in practical settings.

Previous works have introduced diffusion models in AIL for denoising representation learning or loss refinement (notably DiffAIL and DRAIL) but have not utilized the generative capacity of diffusion models to synthesize new expert-like trajectories for direct augmentation of the expert dataset. SD2AIL introduces diffusion-based data synthesis as a core primitive to address expert data scarcity, enabling more effective and sample-efficient adversarial imitation learning.

2. Model Structure and Training Objectives

The SD2AIL algorithm comprises three central modules: (1) diffusion-enhanced discriminator DϕD_\phi, (2) agent policy πθ\pi_\theta learned using Soft Actor-Critic (SAC), and (3) replay buffers for real expert (Re\mathcal{R}_e) and pseudo-expert (Rpe\mathcal{R}_{pe}) samples.

2.1 Notation

  • S,A\mathcal{S}, \mathcal{A}: state and action spaces
  • πe\pi_e: real expert policy / dataset
  • πpe\pi_{pe}: pseudo-expert policy (diffusion-generated, filtered)
  • πθ\pi_\theta: agent policy with parameters πθ\pi_\theta0
  • πθ\pi_\theta1: discriminator output for input πθ\pi_\theta2 with parameters πθ\pi_\theta3
  • πθ\pi_\theta4: total diffusion steps
  • πθ\pi_\theta5: diffusion variances and cumulative products
  • πθ\pi_\theta6: neural network predicting diffusion noise
  • πθ\pi_\theta7: confidence threshold for pseudo-expert filtering
  • πθ\pi_\theta8: mini-batch size of real expert samples
  • πθ\pi_\theta9: pseudo:real sample ratio

2.2 Diffusion Model Loss

The forward diffusion adds noise at each step: DϕD_\phi0

The reverse process is parameterized as: DϕD_\phi1

DϕD_\phi2

The loss for diffusion training is: DϕD_\phi3

2.3 Diffusion-Enhanced Discriminator

The discriminator integrates the diffusion loss as a confidence score: DϕD_\phi4 The surrogate reward for reinforcement learning is: DϕD_\phi5

2.4 Adversarial Objective

The discriminator is trained to output high confidence on both real and pseudo experts and low on agent policy data: DϕD_\phi6

3. Synthetic Demonstration Generation and Filtering

3.1 Reverse Diffusion Sampling

After each discriminator update, pseudo-expert samples are generated by a backward diffusion chain: DϕD_\phi7 yielding trajectory samples DϕD_\phi8.

3.2 Dynamic Confidence-Based Filtering

Samples are admitted to the pseudo-expert buffer only if their discriminator confidence exceeds a dynamic threshold: DϕD_\phi9 This enforces high quality, especially as agent learning progresses.

4. Prioritized Expert Demonstration Replay (PEDR)

PEDR enhances sample efficiency and diversity by prioritizing expert samples (real and pseudo) by their information content as measured by discriminator uncertainty.

  • For each (pseudo-)expert πθ\pi_\theta0, define error πθ\pi_\theta1 and priority πθ\pi_\theta2.
  • Sampling probability:

πθ\pi_\theta3

  • Importance weighting:

πθ\pi_\theta4

  • Discriminator loss with PEDR:

πθ\pi_\theta5

5. Algorithm Details and Implementation

  • Diffusion steps: πθ\pi_\theta6 with linear πθ\pi_\theta7 scheduler
  • Mini-batch: pseudo:real expert ratio πθ\pi_\theta8 (πθ\pi_\theta9)
  • Replay buffers: per-trajectory for experts; pooled for pseudo-experts
  • Networks:
    • Policy Re\mathcal{R}_e0: MLP (256 units × 2 layers, ReLU, Gaussian action heads)
    • Discriminator Re\mathcal{R}_e1: UNet-style encoder + MLP binary classifier
    • Diffusion noise predictor Re\mathcal{R}_e2: shares UNet backbone
  • Optimizer: Adam, learning rate Re\mathcal{R}_e3
  • PEDR parameters: Re\mathcal{R}_e4, Re\mathcal{R}_e5 annealed Re\mathcal{R}_e6
  • Hardware: 3× NVIDIA RTX A6000 GPUs

Pseudo-code (condensed):

  1. Collect agent transitions via Re\mathcal{R}_e7
  2. Sample Re\mathcal{R}_e8 real + Re\mathcal{R}_e9 pseudo-expert samples via PEDR
  3. Compute diffusion and discriminator losses
  4. Update Rpe\mathcal{R}_{pe}0 via gradients
  5. Update PEDR priorities, sample new pseudo-experts, filter by confidence, add to buffer
  6. Compute rewards, update policy via SAC

6. Empirical Results and Analysis

Experiments were conducted on four MuJoCo continuous control tasks (Ant, Hopper, Walker2d, HalfCheetah) with expert datasets of 40 trajectories × 1000 steps, considering low-data settings with 1, 4, or 16 expert trajectories. SD2AIL outperformed baselines such as BC, GAIL, DiffAIL, DRAIL, and SMILING, particularly in 1-trajectory regimes:

Task Expert DiffAIL DRAIL SMILING SD2AIL
Ant 4228 4901 5032 4785 5345
Hopper 3402 3275 3189 3301 3441
Walker2d 5620 5250 5345 5180 5743
HalfCheetah 4663 5600 5720 5501 5885

Ablations showed optimal performance at Rpe\mathcal{R}_{pe}1 diffusion steps, with Fréchet Distance between pseudo and real expert features reduced to 85.4 (compared to 304.7 for random policy) over training. Surrogate reward correlation with true reward achieved 93.0%, 90.1%, 92.3%, 85.2% for SD2AIL across tasks (exceeding DiffAIL). Component ablations on Walker2d established that both pseudo-expert generation and PEDR contribute to final performance (peak return: pseudo-expert only 4557, PEDR only 4907, combined SD2AIL 5743).

7. Discussion, Limitations, and Reproducibility

The principal insight is that diffusion models generate high-diversity, high-fidelity expert-like trajectories, thus addressing the limited support of small real expert datasets and giving the discriminator a more accurate reward boundary. PEDR further prioritizes difficult or uncertain samples, enhancing data efficiency. Limitations include increased wall-clock time due to diffusion sampling; the method remains amenable to acceleration with more efficient diffusion samplers. Empirical results are based on simulated environments; extension to real-world robotics and broader datasets remains open.

SD2AIL is fully reproducible: source code (PyTorch ≥1.10) is provided at https://github.com/positron-lpc/SD2AIL, compatible with MuJoCo and D4RL expert datasets. Standard training is invoked as python train_sd2ail.py --env Hopper --num_traj 1 --T 10 (Li et al., 21 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SD2AIL.