Papers
Topics
Authors
Recent
Search
2000 character limit reached

Return-Conditioned Behavior Cloning

Updated 7 February 2026
  • Return-Conditioned Behavior Cloning is an offline RL method that reframes policy learning as supervised imitation conditioned on return-to-go targets.
  • It simplifies learning by directly imitating demonstrated actions, avoiding the need for value function estimation and enabling stable optimization.
  • Enhancements like ConserWeightive Behavioral Cloning upweight high-return trajectories and apply conservative regularization to address out-of-distribution return challenges.

Return-Conditioned Behavioral Cloning (RCBC) is an offline reinforcement learning (RL) paradigm that recasts policy learning as supervised learning over trajectories, with the key innovation of conditioning the learned policy on a user-specified measure of future return. Rather than inferring value functions or optimizing expected returns through dynamic programming, RCBC directly trains a policy to imitate demonstrated behaviors with additional context in the form of return-to-go (RTG) targets, thereby enabling offline RL with simplified objectives and stable optimization. Extensions such as ConserWeightive Behavioral Cloning (CWBC) address the limitations of naïve return conditioning, particularly in the presence of out-of-distribution (OOD) return requests that require extrapolation beyond the coverage of offline data (Nguyen et al., 2022).

1. Formalism and Training Objective

Given an offline dataset of trajectories D={τ}\mathcal{D}=\{\tau\}, each trajectory τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T), the return-to-go at time tt is defined as

gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.

The return-conditioned policy has the form

πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),

where GG is a user-specified target return, reflecting the intended return level at test time.

Training is conducted via supervised learning by treating each (st,gt)→at(s_t, g_t)\rightarrow a_t tuple as a labeled training example. The typical loss function is the negative log-likelihood of demonstrated actions under the learned policy:

LBC(θ)=−Eτ∼D[∑t=1Tlog⁡πθ(at∣st,gt)].\mathcal{L}_{\rm BC}(\theta) = - \mathbb{E}_{\tau\sim\mathcal{D}}\left[\sum_{t=1}^T\log \pi_\theta(a_t\mid s_t,g_t)\right].

Alternatively, if πθ\pi_\theta is Gaussian and actions are real-valued, a mean squared error can be used:

LBC(θ)=Eτ∼D[1T∑t=1T∥at−πθ(st,gt)∥2].\mathcal{L}_{\rm BC}(\theta) = \mathbb{E}_{\tau\sim\mathcal{D}}\left[\frac1T\sum_{t=1}^T \|a_t-\pi_\theta(s_t,g_t)\|^2\right].

2. Challenges of OOD Return Conditioning

A key deficiency in standard RCBC arises under high-value return conditioning that lies outside the empirical support of the offline dataset. In practice, dataset returns are bounded above by some τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)0; if the policy is queried at test time with τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)1 exceeding this bound, it must extrapolate. This train–test distribution shift in τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)2 pairs often leads to abrupt performance collapse, with realized returns substantially below the requested value. This vulnerability stems from the lack of high-return supervision in training, as well as architectural limitations that hinder generalization to unseen return contexts (Nguyen et al., 2022).

3. ConserWeightive Behavioral Cloning (CWBC): Methodology

CWBC augments RCBC with two principal mechanisms: trajectory weighting and conservative regularization.

a. Trajectory Weighting

Each trajectory τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)3 is assigned a nonnegative weight that increases with total return:

τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)4

with temperature parameter τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)5. This amplifies the influence of high-return trajectories, mitigating the bias toward suboptimal returns endemic to the original offline data distribution.

b. Conservative Regularization

For trajectories with returns exceeding a high percentile τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)6 (e.g., the 95th percentile), pseudo-OOD contexts are constructed by adding noise τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)7 so that τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)8. Define perturbed RTG as

τ=(s1,a1,r1,…,sT,aT,rT)\tau=(s_1,a_1,r_1,\dots,s_T,a_T,r_T)9

The conservative penalty enforced is

tt0

This regularizer constrains the policy’s behavior under OOD returns to remain close to trajectories seen in high-quality data.

c. Combined Objective

The final objective function is:

tt1

where tt2 governs the tradeoff between maximum likelihood imitation and conservative OOD regularization.

4. Implementation Details

The CWBC algorithm involves the following procedural steps (Nguyen et al., 2022):

  • Precompute the returns tt3 for all tt4.
  • Partition trajectories into tt5 return-quantile bins.
  • At each iteration:

    1. Sample a mini-batch: first select a bin tt6 with probability proportional to tt7, then sample uniformly from trajectories in tt8.
    2. For each sampled trajectory, compute the standard BC loss. For those with tt9, inject noise into gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.0 and evaluate the conservative penalty.
    3. Update gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.1 via gradient descent on the combined (weighted + regularized) loss.

Recommended hyperparameters include:

Hyperparameter Value Context
Trajectory-weight bins (gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.2) gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.3 Robust across tasks
Smoothing (gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.4) gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.5 Training stability
Conservative percentile (gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.6) gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.7 OOD regularization
Noise std (gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.8) gt=∑t′=tTrt′ .g_t = \sum_{t'=t}^T r_{t'}\,.9 Return perturbation
Regularization weight πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),0 Final loss function
Test-time conditioning (πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),1) Expert return (no per-task tuning) Evaluation

5. Theoretical Foundations

Appendix C of (Nguyen et al., 2022) provides a bias–variance bound for the gradient discrepancy between the reweighted objective (as used in trajectory weighting) and an ideal expert distribution:

πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),2

where πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),3 is the reweighted return distribution, πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),4 is the number of trajectories with return πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),5, and πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),6 the expert distribution. Exponential weighting is derived as minimizer of this upper bound, balancing bias due to underrepresentation of high-returns and variance introduced by aggressive weighting. Conservative regularization acts as an additional control on extrapolation error for OOD return contexts.

6. Empirical Evaluation

CWBC was evaluated on D4RL locomotion benchmarks (hopper, walker2d, halfcheetah) using “medium,” “med-replay,” and “med-expert” datasets, as well as Atari replay data. Primary metrics included normalized return (relative to expert) and success rates on AntMaze.

Key results (average over 10 seeds):

  • RvS baseline: πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),7 → RvS+CWBC: πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),8 (+18 points)

  • Decision Transformer baseline: πθ(a∣st,gt)or more generallyπθ(a∣st,G),\pi_\theta(a\mid s_t,g_t)\quad\text{or more generally}\quad\pi_\theta(a\mid s_t,G),9 → DT+CWBC: GG0 (+5.2 points)
  • CWBC maintained strong performance even when conditioned on out-of-distribution high returns, often matching or exceeding CQL and IQL.
  • On low-quality “med-replay” datasets, standard RvS crashed on OOD conditioning, whereas RvS+CWBC exhibited consistent reliability.

7. Practical Considerations and Limitations

Empirical evidence supports that CWBC is a robust augmentation to any conditional BC framework, requiring minimal tuning and exhibiting strong generalization for in-distribution and modestly out-of-distribution return conditioning (Nguyen et al., 2022). Robust default hyperparameters further facilitate practical application. However, perfect linear extrapolation to arbitrarily high OOD returns—i.e., guaranteeing proportional performance for any user-specified GG1—remains elusive, and generalization beyond the support of offline data is an unresolved research challenge.

In summary, Return-Conditioned Behavioral Cloning reframes offline RL as supervised learning conditioned on a desired return. ConserWeightive Behavioral Cloning further enhances reliability by (1) upweighting high-return trajectories and (2) imposing a conservative penalty for OOD contexts, collectively yielding a robust and practical recipe for offline RL with minimal complexity (Nguyen et al., 2022).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Return Conditioned Behavior Cloning (RCBC).