Papers
Topics
Authors
Recent
Search
2000 character limit reached

Modality-Composable Diffusion Policy

Updated 8 March 2026
  • MCDP is a diffusion-based policy that composes multiple unimodal diffusion policies via convex score combination during inference.
  • It avoids costly retraining by integrating specialized sensory modalities, such as RGB images and point-clouds, for robust robotic trajectory generation.
  • Empirical evaluations on benchmarks like RoboTwin demonstrate that MCDP improves success rates, offering a modular, plug-and-play approach to multi-modal policy integration.

Modality-Composable Diffusion Policy (MCDP) extends diffusion-based policy models by enabling inference-time composition of multiple pre-trained unimodal diffusion policies, each specialized for a distinct sensor modality. Instead of retraining a single, unified multi-modal policy—a process that incurs significant data and computational cost—MCDP constructs a composite policy by convexly combining the distributional scores (denoising functions) from its constituent unimodal policies during sampling, yielding enhanced adaptability, robustness, and generalization without additional training. Empirical studies in automated robotics tasks, notably on the RoboTwin benchmark, demonstrate that MCDP often outperforms its underlying unimodal policies and establishes a modular, plug-and-play paradigm for integration of arbitrary sensing modalities (Cao et al., 16 Mar 2025, Cao et al., 1 Oct 2025).

1. Theoretical Foundation: Score-Based Diffusion Policies

A diffusion policy (DP) parameterizes a trajectory distribution using a forward noising process and a learned, score-based reverse denoising process. Let τ∈RD\tau \in \mathbb{R}^D denote a trajectory (e.g., robot end-effector pose sequences), and τt\tau_t its noisy counterpart at diffusion step tt. The forward Markov process is defined as: q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right) where αt\alpha_t is a predefined noise schedule. The reverse-time process is parameterized by a neural network ϵθ\epsilon_\theta estimating the noise added at each step: sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t) which provides the score function. The DDPM update for discrete time steps takes the form: τt−1=1αt(τt−1−αt1−αˉtϵθ(τt,t))+σtξ,ξ∼N(0,I)\tau_{t-1} = \frac{1}{\sqrt{\alpha_t}}\left( \tau_t - \frac{1-\alpha_t}{\sqrt{1-\bar{\alpha}_t}}\epsilon_\theta(\tau_t,t) \right) + \sigma_t \xi, \quad \xi \sim \mathcal{N}(0,I) where αˉt=∏i=1tαi\bar{\alpha}_t = \prod_{i=1}^t \alpha_i. In the SDE perspective, the forward SDE is: dτ=−12β(t)τ dt+β(t) dwd\tau = -\frac{1}{2}\beta(t)\tau\,dt + \sqrt{\beta(t)}\,dw and the reverse SDE uses the neural score: τt\tau_t0 Unimodal DPs are pretrained by behavior cloning on trajectory data conditioned on modality-specific inputs:

  • RGB-based DP: τt\tau_t1, trained on RGB images (CNN + transformer encoding).
  • Point-cloud DP: τt\tau_t2, trained on voxelized point clouds.

The objective is the standard diffusion loss: τt\tau_t3 (Cao et al., 16 Mar 2025).

2. Inference-Time Composition: Policy Combination Mechanism

At inference, MCDP convexly combines the scores from τt\tau_t4 pre-trained policies, each operating on a distinct modality τt\tau_t5: τt\tau_t6 where τt\tau_t7 is the noise estimate from the τt\tau_t8-th DP and τt\tau_t9 is its weight. The resulting composite score function is: tt0 In practice, tt1 normalization may be applied to prevent any modality’s score from dominating: tt2

The composite estimate replaces the unimodal score in the reverse denoising step: tt3 Thereby, MCDP generates trajectories under a policy that integrates the strengths of all included modalities (Cao et al., 16 Mar 2025, Cao et al., 1 Oct 2025).

3. Functional Guarantees and System-Level Analysis

The MCDP construction obtains theoretical support in the form of one-step improvement and trajectory-level error bounds (Cao et al., 1 Oct 2025). Given tt4 policies tt5 with time-tt6 score functions tt7, the convexly composed score tt8 remains inside the convex hull of the parent policies: tt9 Define each policy’s mean-squared error to the oracle score q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)0 as q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)1. For q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)2, the error q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)3 of the mixture score q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)4 is convex in q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)5 with minimizer q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)6: q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)7 with strict inequality if the estimators’ errors are non-aligned. This supports the empirical observation that MCDP samplers can outperform all constituent unimodal policies.

For the continuous-time sampling trajectory q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)8 under the composed score, a Grönwall-type bound holds: q(τt∣τt−1)=N(τt;αt τt−1,(1−αt)I)q(\tau_t \mid \tau_{t-1}) = \mathcal{N}\left(\tau_t; \sqrt{\alpha_t}\,\tau_{t-1}, (1-\alpha_t)I\right)9 where αt\alpha_t0, αt\alpha_t1, αt\alpha_t2, αt\alpha_t3, and αt\alpha_t4 are Lipschitz and score-error constants defined in functional analysis of the system-level accuracy (Cao et al., 1 Oct 2025). This suggests that the benefits of composition in single denoising steps can propagate consistently throughout the entire trajectory-generation process.

The standard two-modality MCDP algorithm proceeds as follows:

  1. Initialize with pre-trained unimodal DPs (αt\alpha_t5), their input encodings, and composition weights (αt\alpha_t6).
  2. Sample the noisy trajectory αt\alpha_t7.
  3. For αt\alpha_t8:
    • Compute unimodal noise estimates αt\alpha_t9 and ϵθ\epsilon_\theta0.
    • Linearly blend them: ϵθ\epsilon_\theta1.
    • Update ϵθ\epsilon_\theta2 using the composite ϵθ\epsilon_\theta3.
  4. Return ϵθ\epsilon_\theta4 as the action trajectory.

Grid search on the ϵθ\epsilon_\theta5-simplex is performed over possible weights ϵθ\epsilon_\theta6, using empirical rollout success rates to select optimal ϵθ\epsilon_\theta7. For ϵθ\epsilon_\theta8, a coarse-to-fine grid over ϵθ\epsilon_\theta9 with sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)0 suffices. Each candidate sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)1 is evaluated via multiple rollouts, tracking task success, to select the best-performing mixture for deployment (Cao et al., 16 Mar 2025, Cao et al., 1 Oct 2025).

5. Empirical Evaluation

Quantitative Results: RoboTwin and Robomimic

Empirical tests on the RoboTwin bimanual manipulation suite and Robomimic/PushT show that MCDP generally outperforms both parent unimodal DPs whenever both achieve moderate performance (sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)2 success). For example:

Task DP_img DP_pcd MCDP (best sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)3)
Empty Cup Place 0.42 0.62 0.86 (sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)4)
Dual Bottles Pick (H) 0.49 0.64 0.71 (sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)5)
Shoe Place 0.37 0.36 0.60 (sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)6)
  • If either parent policy is poor (sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)7), MCDP cannot improve over the stronger policy ("Pick Apple Messy").
  • Optimal weights place heavier emphasis on the better-performing DP per task (Cao et al., 16 Mar 2025).

On Robomimic, PushT, and RoboTwin, convex MCDP composition yields sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)8–sθ(τt,t)=−1σtϵθ(τt,t)≈∇τtlog⁡pθ(τt)s_\theta(\tau_t, t) = -\frac{1}{\sigma_t}\epsilon_\theta(\tau_t, t) \approx \nabla_{\tau_t}\log p_\theta(\tau_t)9 improvement on standard benchmarks and approximately τt−1=1αt(τt−1−αt1−αˉtϵθ(τt,t))+σtξ,ξ∼N(0,I)\tau_{t-1} = \frac{1}{\sqrt{\alpha_t}}\left( \tau_t - \frac{1-\alpha_t}{\sqrt{1-\bar{\alpha}_t}}\epsilon_\theta(\tau_t,t) \right) + \sigma_t \xi, \quad \xi \sim \mathcal{N}(0,I)0 in real-world robotic setups:

Method Avg SR Δ vs best parent
DP+MP 41.41 +2.22%
Florence-D+DP 66.76 +5.51%
π₀+FP 88.94 +2.52%

Alternative operators such as logical AND/OR can achieve even greater gains at the cost of per-step recomputation and limited compatibility (e.g., not with flow models) (Cao et al., 1 Oct 2025).

Qualitative Insights

  • Action distributions transition smoothly between the behaviors of each unimodal policy as composition weights are varied, yielding trajectory interpolation.
  • Case studies highlight blending of complementary strengths; e.g., combining approach direction from vision with force estimation from point-cloud data for improved grasping (Cao et al., 16 Mar 2025).

6. Modality and Model Generality

MCDP, via the General Policy Composition (GPC) framework, is agnostic to the sensory modalities and model architectures of its DPs, allowing composition of:

  • Vision-only (RGB), point-cloud, vision–language–action (VLA) policies (e.g., Florence-DiT), and others.
  • Both diffusion- and flow-matching–based policies.

Significant empirical performance gains are observed when parent policies offer complementary strengths. For heterogeneous combinations, MCDP boosts average SR (success rate) by τt−1=1αt(τt−1−αt1−αˉtϵθ(τt,t))+σtξ,ξ∼N(0,I)\tau_{t-1} = \frac{1}{\sqrt{\alpha_t}}\left( \tau_t - \frac{1-\alpha_t}{\sqrt{1-\bar{\alpha}_t}}\epsilon_\theta(\tau_t,t) \right) + \sigma_t \xi, \quad \xi \sim \mathcal{N}(0,I)1–τt−1=1αt(τt−1−αt1−αˉtϵθ(τt,t))+σtξ,ξ∼N(0,I)\tau_{t-1} = \frac{1}{\sqrt{\alpha_t}}\left( \tau_t - \frac{1-\alpha_t}{\sqrt{1-\bar{\alpha}_t}}\epsilon_\theta(\tau_t,t) \right) + \sigma_t \xi, \quad \xi \sim \mathcal{N}(0,I)2 in RoboTwin (vision+point-cloud, VLA+VA pairs), and by similar margins on real-robot tasks (Cao et al., 1 Oct 2025).

Composition requires shared trajectory/action space and diffusion schedule alignment among combined DPs. The approach does not employ classifier-free guidance, thus avoiding doubled computational cost (Cao et al., 16 Mar 2025).

7. Limitations and Extensions

Manual weight tuning is currently required—suboptimal choices can degrade performance, especially if large weights are assigned to poor parent models. MCDP has so far been demonstrated primarily for two visual modalities; extension to additional modalities (e.g., tactile, language-conditioned, proprioceptive DPs) is plausible.

Plausible implications include:

  • Adaptive weight tuning (online or via validation rollouts) could further improve results.
  • Extension to composition across domains and embodiments may be achieved by aligning latent action representations.
  • Investigation of asynchronous modality-specific schedulers and advanced diffusion solvers (such as DPM-Solver or Analytic-DPM) constitutes an open direction (Cao et al., 16 Mar 2025, Cao et al., 1 Oct 2025).

References

  • "Modality-Composable Diffusion Policy via Inference-Time Distribution-level Composition" (Cao et al., 16 Mar 2025)
  • "Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition" (Cao et al., 1 Oct 2025)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Modality-Composable Diffusion Policy (MCDP).