---
title: Dynamic Contrastive Skill Learning
url: https://www.emergentmind.com/topics/dynamic-contrastive-skill-learning-dcsl
type: topic
---

# Dynamic Contrastive Skill Learning

Dynamic Contrastive Skill Learning (DCSL) is a framework for skill discovery and representation in offline reinforcement learning (RL) that integrates state-transition-based skill embeddings, contrastive skill similarity, and adaptive skill-length adjustment. DCSL is designed to resolve limitations of prior skill learning methods—including failure to cluster semantically similar behaviors and rigidity in fixed skill segment lengths—by leveraging contrastive learning and dynamic segmentation. This approach enables flexible skill extraction from complex or noisy demonstrations and improves downstream RL performance on long-horizon, sparse-reward, and noisy-data tasks [2504.14805].

## 1. State-Transition Based Skill Representation

DCSL redefines skill primitives as latent vectors summarizing temporally coherent state transitions instead of fixed-length action blocks. Given an offline dataset $D = \{\tau_i\}_{i=1}^N$ of trajectories $\tau_i = \{(s_t, a_t)\}_{t=1}^T$ where $s_t \in \mathcal{S},\ a_t \in \mathcal{A}$, a skill is a segment starting at time $t$ with (potentially variable) length $H_t$, represented as $z \in \mathcal{Z}$ and capturing state transitions $(s_t \rightarrow s_{t+1}, \dots, s_{t+H_t-1} \rightarrow s_{t+H_t})$.

The embedding process selects four key states per candidate segment: the start $s_t$, two random intermediates $s_{t+a}$ and $s_{t+b}$, and the end $s_{t+H_t-1}$, denoted as $\vec{s}_t = [s_t, s_{t+a}, s_{t+b}, s_{t+H_t-1}]$. An LSTM-based encoder $q_{\theta_q}(z|\vec{s}_t)$ maps this sequence to a skill embedding $z_t$ (where $z_t \sim q_{\theta_q}(z|\vec{s}_t)$). This summarization is regularized through a combination of behavior cloning and prior-matching objectives:

\[
\mathcal{L}_\text{embedding} = \mathbb{E}_{(\vec{s}_t, \vec{a}_t) \sim D} \left[ \mathbb{E}_{z \sim q_{\theta_q}(z|\vec{s}_t)} \left[ \lambda_{BC} \log \pi_{\theta_\pi}(\vec{a}_t | \vec{s}_t, z) - \beta\, KL(q_{\theta_q}(z|\vec{s}_t) \parallel p(z)) + \lambda_{SP} KL(\mathrm{stopgrad}[q_{\theta_q}(z|\vec{s}_t)] \parallel p_{\theta_p}(z|s_t)) \right] \right]
\]

where $\pi_{\theta_\pi}$ is the skill-conditioned action decoder, $p(z) = \tanh(\mathcal{N}(0, I))$ is a prior, and $p_{\theta_p}(z|s_t)$ is a learned skill-prior conditioned on the start state.

## 2. Contrastive Skill Similarity Learning

DCSL introduces an explicit contrastive similarity mechanism to cluster semantically similar skill segments. The similarity function is formulated as

\[
f_{\theta_f}(s, z, s') = \langle \phi_{\theta_\phi}(s, z),\ \psi_{\theta_\psi}(s') \rangle
\]

where $\phi_{\theta_\phi}$ and $\psi_{\theta_\psi}$ are multi-layer perceptrons mapping to a shared $d$-dimensional feature space, with $s$ as the segment start, $z$ the skill embedding, and $s'$ a potential successor state.

For each segment, positive pairs $(s_t, z_t, s_{t+b})$ are constructed where $b$ is a large offset within $[1, H_t-1]$. Negative states $s^-$ are sampled from other skill segments ($z' \neq z_t$). The contrastive (binary) loss is

\[
\mathcal{L}_\text{contrastive} = \lambda_{CL} \mathbb{E}_{(\vec{s}_t, \vec{a}_t) \sim D} \left[ \mathbb{E}_{z \sim q(z|\vec{s}_t)} \left[ -\log \sigma(f_{\theta_f}(s_t, z, s_{t+b})) - \mathbb{E}_{s^-} \log(1 - \sigma(f_{\theta_f}(s_t, z, s^-))) \right] \right]
\]

where $\sigma(\cdot)$ is the logistic sigmoid. This encourages high similarity for a skill’s own successor states and low similarity for states from other segments.

## 3. Dynamic Skill Length Adjustment

Skill length is dynamically determined based on the contrastive similarity function. For a candidate start state $s_t$ and its skill embedding $z_t$, the procedure increments $\alpha$ forward along the trajectory, testing $f_{\theta_f}(s_t, z_t, s_{t+\alpha}) > \epsilon$ for a chosen threshold $\epsilon$ until violation. The resulting length is

\[
H'_t = 1 + \max \{\alpha\ |\ f_{\theta_f}(s_t, z_t, s_{t+\alpha}) > \epsilon\}
\]
with $H'_t$ clamped to interval $[\delta_{min}, \delta_{max}]$. This relabeling procedure is periodically applied to the dataset every $T_\text{relabel}$ steps. The final skill boundaries adaptively reflect the duration over which the skill embedding remains semantically coherent, as judged by the learned similarity.

## 4. Model Training Objective and Algorithm

The overall objective is a weighted sum of the embedding loss, the contrastive loss, and a terminal-state predictor loss:

\[
\mathcal{L}_\text{total} = \mathcal{L}_\text{embedding} + \mathcal{L}_\text{contrastive} + \mathcal{L}_\text{target}
\]

The target loss $\mathcal{L}_\text{target}$ encourages the terminal state of a skill, predicted using the embedding, to align with the observed trajectory outcome via learned encoders and decoders.

Training proceeds by iteratively sampling minibatches, computing all losses, updating all network parameters via Adam (learning rate $3\times 10^{-4}$, batch size 256), and periodically running the skill length relabeling procedure. Key hyperparameters include initial skill length $H_0 = 10$, skill embedding dimension $d_z \in \{2,5\}$, bounds $\delta_{min}=4$, $\delta_{max}=30$, $\epsilon=0$, and loss weights $\lambda_{BC}=2$, $\lambda_{SP}=1$, $\beta=1$, $\lambda_{CL}=1$, $\lambda_{RE}=1$, $\lambda_{ST}=2$.

## 5. Empirical Evaluation and Comparison

DCSL is evaluated across benchmark tasks:

- **AntMaze (D4RL medium-diverse, large-diverse):** Long-horizon navigation with sparse rewards.
- **Kitchen (D4RL mixed-v0):** A complex manipulation task with multiple subtasks.
- **Meta-World Pick-and-Place:** Three settings—expert (ME), medium-replay (MR), full replay (RP, with noise).

Baselines include Behavioral Cloning (BC), Conservative Q-Learning (CQL), CQL+Off-DADS, CQL+OPAL, SPiRL, and SkiMo variants (SkiMo-SAC, SkiMo-CEM).

### Downstream Task Performance (Success Rate)

| Environment  | BC     | CQL       | CQL+Off-DADS | CQL+OPAL | Ours-SAC      |
|--------------|--------|-----------|--------------|----------|---------------|
| AntMaze-M    | 0.0    | 53.7±6.1  | 59.6±2.9     | 81.1±3.1 | 68.0±36.9     |
| AntMaze-L    | 0.0    | 14.9±3.2  | –            | 70.3±2.9 | 73.7±5.9      |
| Kitchen      | 47.5   | 52.4±2.5  | –            | 69.3±2.7 | 94.7±1.5      |

DCSL provides comparable or superior task completion, particularly in the Kitchen task where it significantly outperforms all baselines.

### Sample Efficiency (Timesteps to Success)

| Environment | SPiRL      | SkiMo-CEM  | SkiMo-SAC  | Ours-CEM   | Ours-SAC   |
|-------------|------------|------------|------------|------------|------------|
| AntMaze-M   | 988.5±19.8 | 311.2±95.7 | 833.7±288  | 1000±0     | 453.6±144  |
| AntMaze-L   | 990.2±19.5 | 993.5±13.9 | 881.5±165  | 1000±0     | 672.2±72.9 |
| Kitchen     | 276.6±5.9  | 205.8±29.0 | 251.3±23.7 | 262.0±20.1 | 165.1±4.4  |
| PP (ME)     | 87.8±65.0  | 54.1±21.3  | 58.0±6.8   | 76.0±15.8  | 80.1±13.7  |
| PP (MR)     | 138.0±63.2 | 184.8±24.0 | 87.3±57.6  | 62.9±5.8   | 56.1±5.1   |
| PP (RP)     | 130.6±69.4 | 193.2±13.5 | 200.0±0.0  | 85.1±22.3  | 64.4±16.3  |

Ablations indicate that removing either the contrastive similarity loss or dynamic relabeling degrades robustness, particularly on noisy data.

Skill-space visualizations in AntMaze reveal more diversified and semantically meaningful skill clusters under DCSL compared to prior fixed-length skill VAEs (SPiRL, SkiMo), which tend to collapse into a small repertoire of repetitive patterns. Skill-length distributions inferred by DCSL reflect task structure, with variable-length skills adapting to the environment.

## 6. Distinctive Methodological Contributions

DCSL advances skill discovery and representation by:

- **Skill Embedding via State Transitions:** Encoding skills not as raw action blocks but as latent representations abstracting multi-step state transitions, thus centering the semantic context of behavior.
- **Contrastive Similarity-Based Clustering:** Employing a learned function $f(s, z, s')$ to cluster and differentiate skill segments semantically, informed by contrastive penalties.
- **Dynamic Skill-Length Relabeling:** Regularly re-evaluating segment boundaries based on the learned similarity, yielding context-sensitive skill durations that better align with underlying behavioral motifs.

These innovations enable extraction of more flexible, generalizable, and data-driven skill libraries suitable for hierarchical RL settings and complex offline datasets.

## 7. Context, Limitations, and Implications

DCSL directly addresses limitations in existing fixed-length or VAE-based skill learning by introducing similarity-aware clustering and adaptive temporal abstraction. Results demonstrate improved flexibility on tasks with long horizons, sparse rewards, and imitation from diverse or noisy demonstrations.

A plausible implication is that DCSL’s adaptive mechanism could further benefit multi-task RL or transfer settings where skill distributions and durations are highly variable. However, the approach depends on well-calibrated similarity functions and thresholding. Excessive mismatch between contrastive supervision and actual task semantics could affect discovered skill coherence.

Further extensions could include end-to-end integration with downstream RL, meta-learning of similarity functions, or explicit incorporation of extrinsic task structure. The empirical and methodological contributions of DCSL position it as an advance in unsupervised skill discovery and trajectory abstraction [2504.14805].

Source: https://www.emergentmind.com/topics/dynamic-contrastive-skill-learning-dcsl