Papers
Topics
Authors
Recent
Search
2000 character limit reached

Latent Plan Transformer

Updated 23 December 2025
  • Latent Plan Transformer is a model that introduces a continuous latent plan variable to enable credit assignment and trajectory stitching without step‐wise rewards.
  • It employs a causal Transformer-based trajectory generator conditioned on latent plans, trained via maximum likelihood with MCMC sampling to achieve temporal consistency.
  • Empirical evaluations demonstrate that LPT outperforms baseline models on offline RL benchmarks by ensuring cohesive long-horizon planning and effective trajectory composition.

The Latent Plan Transformer (LPT) is a generative model for trajectory abstraction and planning, specifically designed to address offline reinforcement learning (RL) settings where only full-trajectory returns are available and step-wise reward signals are absent. LPT introduces a latent continuous variable, termed the "plan," which enables the enforcement of temporal consistency across entire episodes, credit assignment over long horizons, and compositional planning via trajectory stitching. Its distinguishing algorithmic feature is planning as latent space inference, realized by a Transformer-based trajectory generator conditioned on the plan variable, with learning and inference achieved via maximum likelihood estimation and Markov Chain Monte Carlo (MCMC) sampling in the latent space (Kong et al., 2024).

1. Problem Formulation and Motivation

LPT is motivated by the challenge of long-term planning with offline RL datasets D={(τi,yi)}D = \{(\tau_i, y_i)\}, where each τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H) is a trajectory of state-action pairs and y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t}) is the total return. In this setting:

  • Credit assignment: Effective association of sparse/delayed rewards to temporally distant actions is difficult in the absence of step-wise reward signals.
  • Trajectory stitching: Construction of new, high-return trajectories from observed suboptimal fragments.
  • Temporal consistency: Mitigation against policy drift in autoregressive models that operate on finite context, conditioned solely on past states and a single summary return.

LPT addresses these issues by introducing a latent "plan" variable zz that generates trajectories and predicts scalar returns, facilitating episode-level coherence and scalable planning as latent variable inference (Kong et al., 2024).

2. Probabilistic Model and Inference

LPT defines a joint generative model: pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma) where

  • pα(z)p_\alpha(z): Plan prior. Implicit, with z0∼N(0,I)z_0 \sim \mathcal{N}(0, I) mapped to z=Uα(z0)z = U_\alpha(z_0) via a neural network (U-Net or MLP).
  • pβ(τ∣z)p_\beta(\tau|z): Trajectory generator. An autoregressive, causal Transformer operating over finite context KK; each token predicts τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)0.
  • τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)1: Return predictor. Gaussian likelihood τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)2, with τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)3 an MLP and τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)4 fixed.

The evidence lower bound (ELBO) for maximum likelihood training is: τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)5 with approximate posterior τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)6. If τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)7, the bound is tight. Marginal likelihood involves integrating out τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)8: τ=(s1,a1,…,sH,aH)\tau = (s_1, a_1, \ldots, s_H, a_H)9 (Kong et al., 2024).

3. Model Architecture and Training

LPT's architecture comprises:

  • Plan prior: Samples y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})0, maps to y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})1, where y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})2 is either a U-Net or MLP, representing an implicit prior y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})3.
  • Trajectory generator: Stack of y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})4 Transformer blocks with causal self-attention over the past y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})5 tokens and cross-attention from y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})6 at each token position. At each timestep y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})7, outputs a Gaussian policy y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})8.
  • Return head: MLP y=∑t=1Hr(s≤t,a≤t)y = \sum_{t=1}^H r(s_{\leq t}, a_{\leq t})9 computing the mean for the Gaussian return predictor.

Training algorithm:

  • LPT is optimized via offline maximum likelihood, leveraging (approximate) posterior sampling in the latent space.
  • For each training example, zz0 is approximated via Langevin dynamics on zz1, where transitions:

zz2

are performed, with gradients over the joint log-probability for the trajectory and return.

  • Model parameters zz3 are updated by gradient ascent using empirical averages over the sampled zz4 (Kong et al., 2024).

4. Planning as Latent Space Inference

At test time, LPT realizes planning as inference by conditioning on a desired return zz5. The plan zz6 is inferred as the mode of zz7 via MCMC in the latent space: zz8 After zz9 steps, set pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)0. An episode is then generated by rolling out the autoregressive policy: pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)1

pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)2

This approach enables specification of arbitrary target returns, framing planning as finding a plan pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)3 most compatible with the desired outcome in the learned latent space (Kong et al., 2024).

5. Empirical Evaluation and Results

LPT is benchmarked on a range of environments:

Domain Tasks/Subsets Characteristics
Gym-Mujoco HalfCheetah, Hopper, Walker2D (medium, replay), AntMaze (umaze, diverse) Continuous control, dense/sparse reward
Maze2D umaze, medium, large Navigation, sparse reward
Connect Four vs. stochastic opponent Board game, adversarial

Baselines: CQL, Decision Transformer (DT), Q-learning Decision Transformer (QDT), Online DT (ODT), ESPER.

Metrics: Average return pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)4 standard deviation over 5 seeds.

Key statistical findings:

  • On Gym-Mujoco with only final return supervision, LPT outperforms DT and QDT, at times matching or exceeding CQL, which has access to step-wise rewards.
  • On Maze2D and AntMaze, LPT yields pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)5–pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)6 improvement over DT by stitching suboptimal trajectory fragments into near-optimal full trajectories.
  • On Connect Four, LPT achieves performance (pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)7) matching SOTA ESPER (pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)8), substantially outperforming DT (pθ(τ,y,z)=pα(z)⋅pβ(τ∣z)⋅pγ(y∣z),θ=(α,β,γ)p_\theta(\tau, y, z) = p_\alpha(z) \cdot p_\beta(\tau|z) \cdot p_\gamma(y|z),\quad \theta = (\alpha, \beta, \gamma)9) (Kong et al., 2024).

Qualitative insights:

  • Posterior gradients in pα(z)p_\alpha(z)0 integrate reward information from sub-trajectories, automating credit assignment.
  • t-SNE visualizations indicate the latent plan space allows interpolation between trajectories, capturing novel, high-return behaviors via trajectory stitching.
  • During execution, the latent pα(z)p_\alpha(z)1 is fixed, but the policy adapts to environment stochasticity, limiting overfitting to dataset contingencies.

6. Strengths, Limitations, and Future Prospects

Strengths:

  • Enforces temporal consistency over entire episodes without explicit step-wise reward or return-to-go conditioning.
  • Posterior sampling in latent space creates abstractions aggregating information across finite-context fragments.
  • Planning as inference (MCMC on pα(z)p_\alpha(z)2) allows for return-conditioned generation without relying on reward-to-go as input.
  • Demonstrates competitive or superior empirical results across dense, sparse, and adversarial benchmarks, supporting long-horizon credit assignment and trajectory composition (Kong et al., 2024).

Limitations and open questions:

  • MCMC-based latent sampling scales poorly for very long horizons. While persistent chains and fewer steps partially mitigate this, further advances such as amortized inference are desirable.
  • The implicit latent prior pα(z)p_\alpha(z)3 lacks explicit density modeling; replacing or augmenting it with normalizing flows or energy-based models may improve expressiveness.
  • Extending LPT to multi-task or hierarchical RL with discrete/discontinuous returns is an open direction.
  • Online continual fine-tuning currently yields limited gains; integrating the LPT posterior machinery with value-based pessimistic objectives remains an open research problem (Kong et al., 2024).
Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Latent Plan Transformer.