Papers
Topics
Authors
Recent
Search
2000 character limit reached

Policy Reconstruction in RL

Updated 31 January 2026
  • Policy Reconstruction is a set of adversarial techniques that use inverse RL to infer unknown reward functions from deployed RL policies.
  • It quantifies reconstruction error using metrics like L1, L2, L∞ distances, demonstrating that standard DP methods may not effectively protect sensitive rewards.
  • Empirical studies reveal that despite tighter privacy budgets, even DP-enhanced RL methods can leak reward details, urging the need for new privacy-centric mechanisms.

Policy reconstruction refers to adversarial techniques for extracting or estimating characteristics of an unknown reward function underlying a published or deployed reinforcement learning (RL) policy. In privacy-sensitive settings such as autonomous driving or recommendation systems, policies trained by RL may implicitly encode private preferences or objectives. "How Private Is Your RL Policy? An Inverse RL Based Analysis Framework" (Prakash et al., 2021) introduces a systematic framework—Privacy-Aware Inverse RL (PRIL)—that formalizes and operationalizes reward-reconstruction attacks against privacy-preserving RL policies. By quantifying the fidelity of reward recovery via inverse RL, the framework exposes substantial gaps between standard differential privacy guarantees applied to policy-training mechanisms and the effective protection of sensitive reward functions.

1. Formalization of the Reward-Reconstruction Attack

The PRIL framework defines the reward-reconstruction attack as an adversarial process: given a deployed RL policy π\pi trained on an unknown reward function R:S→RR : S \to \mathbb{R}, an adversary applies inverse RL to infer a reconstructed reward R^\hat R. The framework evaluates both non-private policies (π′\pi') and differentially-private policies (π′′\pi''), measuring privacy via a set of reward-distance metrics between RR and the recovered R^\hat R:

  • Non-private baseline: π′\pi' trained directly on RR.
  • Private policy: π′′\pi'' trained on R:S→RR : S \to \mathbb{R}0 with a chosen differential privacy (DP) mechanism.

The attack pipeline is:

  1. Apply inverse RL to R:S→RR : S \to \mathbb{R}1 to obtain R:S→RR : S \to \mathbb{R}2.
  2. Apply inverse RL to R:S→RR : S \to \mathbb{R}3 to obtain R:S→RR : S \to \mathbb{R}4.
  3. Compute distances R:S→RR : S \to \mathbb{R}5 and R:S→RR : S \to \mathbb{R}6 using several metrics.
  4. If R:S→RR : S \to \mathbb{R}7 is large, the private policy R:S→RR : S \to \mathbb{R}8 offers strong reward privacy. Otherwise, R:S→RR : S \to \mathbb{R}9 is vulnerable to reconstruction.

2. Inverse RL Algorithm: Finite-State LP Formulation

Reward reconstruction leverages the classical finite-state Inverse RL algorithm of Ng and Russell (2000), instantiated as a linear program (LP) that finds the minimum-norm reward vector R^\hat R0 making the input policy R^\hat R1 uniquely optimal by at least a margin of 1. The optimization is:

  • Objective:

R^\hat R2

  • Optimality-by-margin constraints (for every state R^\hat R3 and every action R^\hat R4):

R^\hat R5

where

R^\hat R6

R^\hat R7

The full LP can be written as:

R^\hat R8

The LP returns R^\hat R9 as the reconstructed reward vector consistent with policy π′\pi'0.

3. Differential Privacy in RL Algorithms

Three major classes of DP mechanisms are evaluated in PRIL:

  • Value Iteration + DP-Bellman (VI-DP-Bellman):
    • Gaussian noise π′\pi'1 is added to each Bellman update:

    π′\pi'2 - Sensitivity for the update is π′\pi'3. - Rényi-DP theory dictates:

    π′\pi'4

  • Deep Q Network (DQN): DP-SGD, DP-Adam, DP-Shoe, DP-FN:

    • DP-SGD/Adam/Shoe: per-example gradients are clipped to norm π′\pi'5 and Gaussian noise π′\pi'6 is added.
    • DP-Shoe uses SGD + tanh activations; DP-SGD/DP-Adam use ReLU.
    • DP-FN: functional Gaussian process noise is added directly to Q-value estimates, ensuring π′\pi'7-DP.
  • Proximal Policy Optimization (PPO):
    • Only the actor network updates are privatized using DP-SGD, DP-Adam, or DP-Shoe.
    • Critic updates remain non-private.
    • The same gradient clipping and Gaussian noise are applied as in DQN.

Privacy parameters in experiments:

  • π′\pi'8
  • π′\pi'9
  • π′′\pi''0 is computed using TensorFlow-Privacy’s RDP-to-DP conversion.

4. Quantitative Reward-Distance Metrics

PRIL assesses reconstruction error with four metrics, operating on reward vectors π′′\pi''1 and π′′\pi''2 (normalized):

Metric Formula Description
π′′\pi''3 distance π′′\pi''4 Total variation
π′′\pi''5 distance π′′\pi''6 Euclidean error
π′′\pi''7 distance π′′\pi''8 Maximum deviation
Sign-change count π′′\pi''9 Reward sign flips across states

These metrics directly quantify the adversary’s ability to reconstruct the true underlying reward.

5. Empirical Evaluation: FrozenLake Benchmarks and Results

Experiments are carried out on 24 customized FrozenLake domains (12 of RR0 and 12 of RR1 grids), featuring cell types: Safe (S), Frozen (F), Hole (H), High-reward (A), and Goal (G), with near-deterministic transitions (slip factor RR2).

  • Algorithms Evaluated: VI-DP-Bellman, DQN-DP-SGD, DQN-DP-Shoe, DQN-DP-Adam, DQN-DP-FN, PPO-DP-SGD, PPO-DP-Shoe, PPO-DP-Adam, plus non-private baselines.
  • Training Protocols: DQN/PPO use 15 epochs, 200 iterations, batch size 50, micro-batches 5, learning rate 0.15, discount RR3. VI runs Bellman updates until convergence.

Policy extraction and reward reconstruction are performed for each RR4 and algorithm over 10 random seeds. Distances RR5, RR6, RR7, and RR8 are computed.

Key findings:

  • All four reconstruction error metrics are flat as a function of RR9: increasing the privacy budget does not substantially decrease R^\hat R0.
  • Aggregated R^\hat R1 error ranks: DQN-variants R^\hat R2 PPO-variants R^\hat R3 VI-DP-Bellman. E.g., R^\hat R4 (DQN), R^\hat R5 (PPO), R^\hat R6 (VI-DP-Bellman) for R^\hat R7 environments.
  • The utility–privacy trade-off is present for policy returns (lower R^\hat R8 reduces return), but reward reconstruction error remains insensitive to R^\hat R9.

6. Privacy-Gap Analysis and Implications

A fundamental insight is that differential privacy mechanisms protecting policy updates—such as per-iteration gradient noise or DP-Bellman—do not guarantee privacy for the underlying reward function. Inverse RL attacks in PRIL can reconstruct π′\pi'0 from π′\pi'1 with a small, constant error substantially independent of π′\pi'2. Classical non-deep methods (VI-DP-Bellman) provide less reward privacy than deep RL approaches (DQN/PPO), yet the improvement is insufficient for privacy-critical domains.

This suggests a significant mismatch between the theoretical privacy guarantees (in terms of π′\pi'3-DP applied during training) and the protection actually required to conceal the true reward. A plausible implication is that future DP-RL research must either develop reward-centric privacy mechanisms—directly privatizing π′\pi'4—or redefine DP guarantees to explicitly bound reward-reconstruction error under adversarial inverse RL.

7. Summary and Directions

PRIL provides a rigorous framework for adversarial assessment of reward privacy in RL. It leverages finite-state LP inverse RL to reconstruct the reward from policies trained under various DP mechanisms and quantifies the privacy leakage via multiple reward-distance metrics. Empirical results clearly indicate a gap: prevailing DP-RL techniques do not prevent effective reward reconstruction, exposing the need for fundamentally stronger privacy definitions and mechanisms in RL policy deployment (Prakash et al., 2021).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Policy Reconstruction.