DeLock: Advanced Locking and VLA Methods
- DeLock is a dual mechanism that includes a decoupled locking protocol for disaggregated memory and a method to mitigate lock-in in vision-language-action policies.
- Its memory protocol decouples state management from ownership transfer using one-sided RDMA atomics and direct CN-to-CN notifications, reducing NIC contention and boosting throughput.
- Its VLA approach preserves visual grounding through drift regularization and contrastive prompt guidance, significantly improving out-of-distribution performance.
DeLock refers to two distinct advanced mechanisms in computer systems research: (1) a locking protocol for disaggregated memory architectures designed to mitigate network interface controller (NIC) contention by decoupling state maintenance and ownership transfer (Zhang et al., 23 May 2025), and (2) a method for mitigating "lock-in"—the loss of generalization—in vision-language-action (VLA) policies after low-data supervised fine-tuning, by preserving visual grounding and applying contrastive prompt guidance (Huang et al., 25 Apr 2026). Both innovations address challenges of steerability and resource bottlenecks in large-scale distributed or data-driven systems.
1. DeLock for Disaggregated Memory: System Model and Bottleneck Analysis
Disaggregated memory (DM) architectures consist of CPU-rich compute nodes (CNs) with minimal local DRAM and memory-rich memory nodes (MNs) with large DRAM pools but limited CPU, interconnected via high-speed RDMA-capable networks. In such systems, locks for protecting hot data are typically placed on MNs. Traditional RDMA-based spinlocks or queueing locks (e.g., MCS, ShiftLock) impose significant NIC load due to repeated remote atomic operations or reader polling, especially under high contention. Empirical analysis shows that with 256 clients, RDMA spinlocks cause retry storms, consuming over 95% of available MN-NIC IOPS, reducing application throughput to below 5% of ideal and increasing 99th-percentile latency up to 69× (Zhang et al., 23 May 2025).
2. Decoupled Locking Design: Cooperative Queue-Notify Protocol
DeLock’s approach for DM separates centralized state management from decentralized ownership coordination. The lock state—mode (exclusive/shared), metadata, and the queue itself—resides in a compact, fixed-size structure on the MN. Ownership transfer is realized through direct CN-to-CN notifications, avoiding unnecessary round-trips through the MN. The protocol relies on a 64-bit control header encoding queue head, size, waiters, and reset signals, atomically updated using RDMA fetch-and-add (FAA). Queue entries (8 B each) store mode, client ID, version, and optionally timestamp for hierarchical fairness.
During acquisition, a CN attempts to increment the queue; if contention is detected, it writes to its queue entry and blocks awaiting a notification. Release involves a single FAA decrement and read, followed by a minimal number of direct CN-to-CN notifications to transfer lock ownership. This minimizes MN-NIC traffic and bounds NIC IOPS per operation to , regardless of system contention.
3. Fairness, Queueing, and State Transitions
DeLock implements strict FIFO ordering by queue index and supports both exclusive and shared (reader) entries. When an exclusive holder releases, subsequent shared waiters are discovered through the queue and notified en masse. Hierarchical extensions using per-entry timestamps enable fair ordering across logical CN boundaries or datacenter phases. The state machine cycles between FREE, HELD, and WAITING states, driven by atomic FAA transitions and ownership transfers mediated outside the MN.
4. Implementation and Comparative Performance
DeLock is realized using one-sided RDMA atomics (FAA, READ, WRITE) directed at the MN and reliable two-sided RDMA SENDs between CNs for notifications. Queueing metadata fits in an 8 B header plus queue entries, with chosen larger than the number of CNs to avoid slow, rare resets. Evaluations show up to 43.37× throughput improvement over RDMA spinlocks and 1.81× over MCS locks in microbenchmarks with skewed (Zipf ) workloads. Case studies with a disaggregated object store and the Sherman B⁺Tree index show, for example, a 35.60× object store throughput increase over CASLock and 2.31× over cohort locks, with 99th percentile latency reductions up to 98.8% (Zhang et al., 23 May 2025).
| Application | Baseline | Throughput Ratio (DeLock/baseline) | 99th-pct Latency Improvement |
|---|---|---|---|
| Object Store | CASLock | 35.60× | 98.8% |
| B⁺Tree Index | Cohort Lock | 2.31× | 82.1% |
Trade-offs include a dependency on efficient CN-to-CN messaging (15% throughput loss observed if CN-CN RTT exceeds CN-MN RTT), the need for infrequent time synchronization, and configuration of queue capacity to avoid resets.
5. DeLock for Vision-Language-Action Policies: Motivation and Method
In the VLA context, DeLock addresses "lock-in," where low-data fine-tuning on demonstrations causes a generalist policy to overfit to seen instructions (concept lock-in) or spatial targets (spatial lock-in), effectively disabling generalization to novel tasks. For instance, a VLA policy may output identical actions for unseen prompts with different target objects or locations, if those distinctions were absent in the fine-tuning data. This phenomenon arises even when the pre-trained policy was originally broad in scope (Huang et al., 25 Apr 2026).
6. Post-Training Protocol: Visual-Drift Regularization and Prompt Guidance
The DeLock method employs a pre-trained VLA policy with frozen base weights in the language backbone and action head, augmented by LoRA adapters inserted for downstream adaptation. The core innovation is a visual-encoder drift regularizer: the post-training loss is composed of behavioral cloning on the small demonstration set combined with an penalty keeping the visual encoder close to its pre-trained weights:
This constraint preserves the model’s capacity to visually ground novel concepts and spatial relations, mitigating lock-in. Inference-time contrastive prompt guidance (CPG) is also introduced: at every rollout step, the predicted denoising field for a positive prompt is contrasted against a negative prompt , and the update is steered via a scaled difference, leading to improved alignment with novel instructions without retraining.
7. Empirical Results and Implications
On eight simulation and real-world tasks probing out-of-distribution (OOD) location and prompt shifts, DeLock achieves superior success rates on novel instructions compared to baselines, including high-data retrained policies. For instance, on fine-grained spatial shifts, DeLock outperforms RETAIN and π₀.₅-DROID, and ablation studies indicate that both visual-drift regularization and CPG are necessary for OOD generalization—omitting either substantially degrades performance.
| Method | T5 [Spatial] | T6 [Spatial] | T8 [Concept+Spatial] |
|---|---|---|---|
| RETAIN | 0/20 | 0/20 | 1/20 |
| π₀.₅-DROID | – | – | 0/20 |
| DeLock (full) | 11/20 | 13/20 | 13/20 |
| DeLock w/o CPG | 0/20 | 0/20 | 0/20 |
| DeLock w/o Vis-Reg | 0/20 | 0/20 | 0/20 |
A plausible implication is that regularizing the visual encoder during SFT is critical for generalist robot policies in low-data regimes, and CPG-style steering can recover generalization not achieved by SFT alone (Huang et al., 25 Apr 2026).
References
- "DecLock: A Case of Decoupled Locking for Disaggregated Memory" (Zhang et al., 23 May 2025)
- "Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training" (Huang et al., 25 Apr 2026)