Papers
Topics
Authors
Recent
Search
2000 character limit reached

Remember what you did?: Learning Behavioral Memories for Partially Observable Object Manipulation

Published 19 Jun 2026 in cs.RO | (2606.21188v1)

Abstract: Long horizon, contact-rich manipulation is inherently partially observable. This is as a single visual observation rarely captures a robot's full action context, including prior attempts, interactions, or progress. Consequently, standard visuomotor policies or vision-language-action models are prone to struggle in such tasks due to a lack of memory. To address this, we introduce Compressed Action Memory Policy (CAMP) based on the insight that a robot's own action history serves as a highly informative, self-supervised signal, enabling the policy to learn a robust, compact history representation. In our approach, we train a memory module to maintain a compressed representation of past actions, forcing it to encode a latent behavioral memory of all the robot's past interactions that can then be used to better contextualize future actions. This allows our approach to implicitly track generalized task progress and learn from failed attempts without any additional supervision, or external oversight. We evaluate CAMP across four real-robot setups and two novel simulation benchmarks: Memory-T-Bench and Memory-Manip-Bench. By demonstrating substantial gains over state-of-the-art baselines, CAMP is, to our knowledge, the first policy to demonstrate substantial success on contact-rich partially observable manipulation tasks purely through learned memory.

Summary

  • The paper introduces CAMP, a policy that uses compressed action memory to overcome partial observability in contact-rich robotic manipulation.
  • It employs DCT-based trajectory compression, a two-layer LSTM, and vector quantization to distill behavioral cues essential for multi-stage tasks.
  • Empirical results highlight up to 70% improvement in success rates over memoryless baselines across simulation and real-robot benchmarks.

Compressed Action Memory Policy (CAMP): Learned Behavioral Memory for Contact-Rich Manipulation Under Partial Observability

Motivation and Context

Partial observability poses a central obstacle in long-horizon, contact-rich robotic manipulation tasks. Typical visuomotor policies, including vision-language-action (VLA) models, rely heavily on current or short-term observation histories, often failing to encode critical behavioral context such as prior interactions or failed attempts. Existing memory-augmented approaches predominantly focus on visual or semantic memory, which are insufficient for capturing the non-visual, action-centric aspects of contact-rich manipulations. This paper introduces CAMP, a policy architecture that explicitly leverages the robot's action history to improve decision-making in partially observable, multi-step tasks, bypassing the need for high-dimensional visual sequence prediction or costly manual prompting.

Methodology

CAMP replaces reliance on visual or language-prompted histories with a memory module trained to reconstruct compressed representations of past action sequences. The recurrent module, instantiated as a two-layer LSTM, encodes each observation-action pair stepwise, combining visual features from a fixed third-person camera with the preceding proprioceptive and action states.

The core technique is to compress the past trajectory using the Discrete Cosine Transform (DCT), which efficiently encapsulates the low-frequency structure of actions, discarding irrelevant high-frequency details. Pretraining aims for frequency-weighted reconstruction and temporal consistency: the module penalizes disagreement across prediction offsets to enforce stable behavioral memory over time. The memory is further discretized via vector quantization, projecting hidden states onto a learned codebook, improving generalization and preventing overfitting to demonstration-specific trajectories.

The action module is a diffusion-based policy, conditioned not only on current visual-proprioceptive input but also on the compact memory code, generating action chunks by iteratively denoising Gaussian samples. Training follows a staged regime: first freezing the memory module and allowing the action head to adapt, then jointly finetuning, enabling the memory representation to align with task-specific action distributions.

Empirical Evaluation

The authors introduce two novel simulation benchmarks—Memory-T-Bench (2D tasks) and Memory-Manip-Bench (3D tasks)—explicitly designed to require memory of past actions for success, not merely visual state reconstruction. Additionally, CAMP is validated across four challenging real-robot tasks using a Franka Emika Panda arm and dual RGB streams.

On Memory-T-Bench, CAMP achieves success rates of up to 94%–98% across tasks, in stark contrast to memoryless baselines (e.g., Diffusion Policy), which drop to as low as 48% under severe partial observability. Critically, the performance gap widens precisely in task variants where visual observations conceal task progress or prior failures, indicating robust disambiguation of hidden state via action memory.

In 3D tasks, CAMP maintains a strong margin, averaging 64.3% success on Memory-Manip-Bench versus 40.9% for the best memoryless baseline. Notably, in tasks demanding explicit recall of prior actions (Swap-Block: 86% vs. 18%; Uncover-Blocks: 36% vs. 6%), CAMP consistently outperforms Memory VLA and To.5, including those with large-scale pretraining or visual-token memory. On real-robot tasks, all baselines collapse to zero or near-zero success except CAMP, which shows non-trivial rates across each scenario, revealing profound practical implications for embodied decision making.

Ablation studies demonstrate that CAMP's training schedule, vector quantization, and DCT-based compression are all essential for optimal performance. Removing discretization reduces transfer across demonstrations, and departing from DCT (e.g., using Fourier or raw waypoints) also degrades results. The reconstruction loss installed by pretraining serves mainly as a scaffold; finetuning the entire policy causes the memory module to move away from raw trajectory reconstruction, instead optimizing toward compressed behavioral summaries aligned with action selection.

Claims and Numerical Results

  • CAMP achieves up to 70% absolute improvement in success rate on tasks where partial observability precludes single-frame policies.
  • Performance margins scale directly with the degree of hidden state: CAMP imposes no penalty on fully observable problems while supplying decisive advantage where state aliasing impedes traditional approaches.
  • The memory module representation, after finetuning, no longer reconstructs the exact trajectory but instead distills key behavioral cues for decision making, highlighting that optimal memory for action policy diverges from optimal trajectory reconstruction.
  • CAMP is the only method to achieve consistent completion of multi-stage memory-dependent tasks in real-world evaluations.

Implications and Future Directions

CAMP fundamentally reframes memory-augmented visuomotor policy learning by exploiting action history as a compact, self-supervised learning signal. This approach is particularly consequential in contact-rich settings where visual observations alone are insufficient, and where explicit recall of prior actions or failures is essential for task progress.

Practically, CAMP enables more scalable and autonomous robotic manipulation, eliminating the dependence on handcrafted prompts, high-dimensional visual retrieval, and expensive annotation. Theoretically, this work demonstrates that behavioral memory formation through compressed action sequence prediction resolves long-standing bottlenecks in memory-dependent policy learning under partial observability.

Future research directions include extending CAMP to tasks with extreme-horizon requirements and generalizing the approach to dexterous, multi-object manipulations. Further exploration of alternative memory compression schemes and their influence on policy generalization may refine the trade-off between reconstructive fidelity and behavioral abstraction.

Conclusion

The paper presents CAMP, a memory-augmented visuomotor policy leveraging compressed action history to resolve partial observability in contact-rich manipulation tasks. Through rigorous empirical evaluation, including challenging benchmarks and real-robot experiments, CAMP demonstrates clear superiority over baselines that rely on observation history or visual memory alone. The findings establish compressed behavioral memory as a critical enabler for autonomous robotic manipulation, opening rich avenues for further exploration in embodied AI (2606.21188).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.