Papers
Topics
Authors
Recent
Search
2000 character limit reached

AtomSkill Framework for Robotic Manipulation

Updated 3 January 2026
  • AtomSkill is a multi-task imitation learning framework that uses semantically grounded, variable-length atomic skills to address noisy, multi-modal robotic demonstrations.
  • It employs contrastive clustering and vision–language annotation for precise segmentation and semantic alignment of skills from demonstration trajectories.
  • A diffusion-based keypose imagination module enables robust long-horizon planning and composable action generation for versatile robotic manipulation tasks.

AtomSkill is a multi-task imitation learning framework developed to address the challenges of scalable robot manipulation across diverse tasks. Conventional imitation learning approaches excel in single-task domains but encounter performance degradation in multi-task settings due to demonstration noise, behavioral multi-modality, suboptimal skill segmentations, and limited abstraction for long-horizon planning. AtomSkill advances the domain by learning a structured Atomic Skill Space, enabling compositionality and semantic skill reuse without reliance on fixed-length segmentation or environment-specific priors. Its core contributions include the discovery of semantically grounded variable-length atomic skills, contrastive clustering for temporal and semantic coherence, and a Keypose Imagination module for robust chaining and action generation (Zhu et al., 20 Dec 2025).

1. Framework Architecture and Motivation

AtomSkill processes robot demonstrations comprising observation-action sequences and natural language instructions. The architecture is motivated by three challenges: noisy and multi-modal demonstrations, ambiguity from fixed-length skill segmentation, and a lack of long-horizon planning abstractions in prior skill-based methods. To address these issues, key components of the workflow are:

  • Skill segmentation: Demonstrations are partitioned into non-overlapping segments {τ1,…,τn}\{\tau_1,\dots,\tau_n\} where binary gripper state changes, ensuring segments align with meaningful contact events and atomic skill boundaries.
  • Semantic annotation: A large pre-trained vision–LLM provides natural-language labels (e.g., “grasp,” “place”) for each segment, facilitating semantic grounding.
  • Skill encoding and compression: Variable-length action sequences are encoded by a VQ-VAE-style latent encoder ϕθ\phi_\theta into nn latent tokens, which are quantized into a discrete codebook E={ek}k=1KE=\{e_k\}_{k=1}^K.
  • Skill sampling for planning: A diffusion-based prior ρθ\rho_\theta enables high-level skill planning by sampling from the atomic skill space during inference.
  • Action decoding with keypose imagination: The action decoder ψθ\psi_\theta incorporates current observations, skill embedding, and a “keypose” token to jointly predict short-term action sequences and long-horizon terminal keyposes.

This design facilitates efficient skill chaining and composability, with skill boundaries and transitions mediated by spatial proximity to predicted keyposes, rather than handcrafted heuristics.

2. Semantically Grounded Atomic Skill Library

The construction of a skill library within AtomSkill involves a multi-stage pipeline:

2.1 Gripper-State Keyframe Detection

Demonstration trajectory τ={(Ot,at)}t=1T\tau = \{(O_t, a_t)\}_{t=1}^T with instruction LL is segmented at time points where the gripper state flips (open/close), yielding variable-length segments τi\tau_i. This mechanism aligns abstract skill representations with interaction events, supporting robust skill discovery.

2.2 Vision–Language Annotation

Each segment τi\tau_i is annotated by querying a vision–LLM (e.g., Qwen2–VL) with segment images and the global instruction ϕθ\phi_\theta0. The model generates segment descriptions ϕθ\phi_\theta1 and discrete semantic labels ϕθ\phi_\theta2, yielding a labeled skill dataset ϕθ\phi_\theta3.

2.3 Contrastive Clustering in Skill Embedding Space

Continuous skill embeddings ϕθ\phi_\theta4 are quantized into the codebook. To enforce the desired properties:

ϕθ\phi_\theta5

  • Supervised contrastive objectives: Embeddings are clustered with respect to both intra-skill temporal coherence and inter-task semantic alignment:
    • Temporal contrastive loss ϕθ\phi_\theta6 for consistency across token positions within a skill.
    • Semantic contrastive loss ϕθ\phi_\theta7 for clustering across tasks sharing semantic labels.
    • Total contrastive loss: ϕθ\phi_\theta8.

This dual-objective structure yields a compact, discrete codebook facilitating generalization and skill reuse across tasks.

3. Action Generation and Skill Chaining via Keypose Imagination

The action decoder ϕθ\phi_\theta9 is designed to handle multi-modal inputs:

  • Inputs: Multi-view images nn0, proprioceptive state nn1, task instruction nn2, and the current discrete skill token sequence nn3.
  • Architecture: Cross-attention modules and temporal self-attention enhance context-integration and sequential dependency modeling.
  • Outputs:
    • Action chunk head: Predicts actions nn4 for an action horizon nn5, with reconstruction loss nn6.
    • Keypose head: Predicts terminal action nn7, providing long-horizon intent, trained with keypose loss nn8.

During inference, skill embeddings sampled from nn9 determine the next skill to execute, and action chunk execution proceeds until the predicted action is within E={ek}k=1KE=\{e_k\}_{k=1}^K0 of the keypose, capturing the natural skill transition boundary.

4. Training and Inference Algorithms

4.1 Training Procedure

AtomSkill training proceeds in two stages according to the following algorithmic outline:

ρθ\rho_\theta7

Full objective:

E={ek}k=1KE=\{e_k\}_{k=1}^K1

4.2 Inference Workflow

Inference proceeds in iterative skill selection and execution:

ρθ\rho_\theta8

Skill transitions are governed by proximity of executed actions to predicted keyposes, supporting robust chaining.

5. Implementation Specifications and Experimental Results

5.1 Network Design and Hyperparameters

  • Encoder E={ek}k=1KE=\{e_k\}_{k=1}^K2: 1D-CNN layers with 6-layer self-attention, output E={ek}k=1KE=\{e_k\}_{k=1}^K3 tokens.
  • Codebook: E={ek}k=1KE=\{e_k\}_{k=1}^K4 entries, commitment E={ek}k=1KE=\{e_k\}_{k=1}^K5.
  • Decoder E={ek}k=1KE=\{e_k\}_{k=1}^K6: Cross-attention module with 7 layers, 8 heads; action horizon E={ek}k=1KE=\{e_k\}_{k=1}^K7.
  • Diffusion sampler E={ek}k=1KE=\{e_k\}_{k=1}^K8: CNN-based U-Net with FiLM conditioning.
  • Optimization: Learning rates E={ek}k=1KE=\{e_k\}_{k=1}^K9, weight decay ρθ\rho_\theta0, batch size ρθ\rho_\theta1, loss weights ρθ\rho_\theta2, ρθ\rho_\theta3, ρθ\rho_\theta4.

5.2 Empirical Evaluation

AtomSkill demonstrates superior quantitative performance in multi-task robotic manipulation:

Setting ATP SR (%) Comparison Benchmarks
RLBench (6 tasks) 0.68 67.2 DP: 0.54/37.2, ACT: 0.55/46.7, VQ-BeT: 0.10/5.0, QueST: 0.39/30.0
Real-world bimanual (300 demos) 0.60 — ACT: 0.34, RDT: 0.28

Ablation studies reveal the necessity of contrastive losses (ATP drops to 0.33 without ρθ\rho_\theta5 and ρθ\rho_\theta6), and keypose imagination yields significant performance gains, especially on spatially localized tasks (ATP up from 0.61 to 0.68, SR up from 53.9% to 67.2%).

This suggests that semantically grounded, temporally coherent atomic skills coupled with keypose-conditioned action decoding materially improve composability and robustness in multi-task manipulation.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AtomSkill Framework.