Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bundled Contact Gradients (BCG): A Gradient Stabilization Method

Updated 29 September 2026
  • Bundled Contact Gradients (BCGs) are a contact-local randomized-smoothing technique for differentiable policy learning, used to stabilize gradients by evaluating multiple nearby perturbation rollouts and averaging their state sensitivities.
  • BCGs are activated during stiff contacts and utilize perturbations in Cartesian space to adaptively stabilize policy gradients in dynamic humanoid systems, enhancing simulation fidelity and learning efficiency.
  • Experimental evaluations on a Unitree G1 humanoid robot across various tasks like running, jumping, and fighting show that BCG reduces gradient variance and improves transfer reliability to a greater extent than SHAC and PPO baseline models, although with higher computational overhead.

Bundled Contact Gradients (BCG) are a contact-local randomized-smoothing method for differentiable policy learning introduced to stabilize gradients through stiff rigid-body contact. When a stiff contact is detected, BCG evaluates multiple nearby perturbation rollouts around the contacting configuration and averages their state sensitivities before continuing the nominal trajectory. The forward simulator retains finite, stiff contact dynamics, while the gradient signal is locally smoothed to reduce variance caused by contact timing, impact sensitivity, and rapidly changing transition Jacobians (Aditya et al., 25 Sep 2026). The acronym BCG is context-dependent: it denotes “Bundled Contact Gradients” in differentiable simulation, but “bundled gradient” has a related meaning in randomized smoothing for contact-aware optimal control (Suh et al., 2021), while other literature uses BCG for ballistocardiography and brightest cluster galaxy.

1. Terminology and conceptual scope

Bundled Contact Gradients concern the estimation of derivatives through contact-sensitive dynamical systems. They should be distinguished from several unrelated uses of the acronym:

  • Bundled Contact Gradients: a contact-local randomized-smoothing technique for differentiable policy learning (Aditya et al., 25 Sep 2026).
  • Bundled gradients in contact optimal control: smoothed gradients or Jacobians obtained by averaging derivatives over perturbed state–input points, used in iterative randomized-smoothing MPC (Suh et al., 2021).
  • Ballistocardiography: an ambient mechanical biosignal measured from cardiac recoil, used in the BCG-FM foundation model (Kjaer et al., 5 Jun 2026).
  • Brightest cluster galaxy: the central luminous galaxy in a galaxy cluster, abbreviated BCG in observational astronomy (Alcorn et al., 2023, Zenteno et al., 27 Mar 2025).

In the differentiable-simulation formulation, the underlying transition is

xt+1=fκ(xt,at),x_{t+1}=f_\kappa(x_t,a_t),

where xtx_t is the simulator state, ata_t is the action, and κ\kappa is the contact stiffness. The policy gradient depends on the state sensitivity

∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.

BCG modifies the derivative evaluation near detected contact events rather than replacing the policy, the forward contact law, or the policy-optimization objective.

The term “bundled” refers to aggregation across a local collection of contact realizations. It does not denote a new physical contact law, a conventional gradient of the nominal trajectory, or a global smoothing operation applied indiscriminately to the simulator.

2. Motivation: stiff contact and unstable sensitivities

Differentiable simulation provides analytic derivatives of robot dynamics and can be substantially more sample-efficient than model-free policy-gradient estimators. However, rigid-body contact produces a difficult trade-off between physical fidelity and gradient regularity.

The differentiable simulator considered in the BCG formulation uses an analytically smoothed contact activation,

sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},

where d>0d>0 denotes penetration and d<0d<0 denotes separation. Small κ\kappa produces broad, compliant contact; large finite κ\kappa produces sharp, stiff contact; and the limit xtx_t0 approaches hard-contact-like behavior.

Soft contact is easier to differentiate, but excessive compliance permits unrealistic penetration and delayed or weak force buildup. Policies trained under such dynamics may learn behaviors that do not transfer to hardware. Increasing xtx_t1 improves contact fidelity but makes the transition Jacobians highly sensitive to small perturbations in position and velocity near contact. Nearby contact configurations can then yield sharply different derivatives, resulting in high-variance policy updates, exploding or poorly conditioned sensitivities, dependence on individual contact realizations, and failure to learn dynamic behaviors despite physically appropriate forward dynamics (Aditya et al., 25 Sep 2026).

The issue is especially consequential for dynamic humanoid motions. Running, jumping, fighting, and dancing depend on foot-ground impulses, impact timing, balance recovery, and whole-body momentum transfer. A compliant model can distort those interactions, whereas a stiff model can make direct backpropagation unreliable.

BCG addresses this trade-off by retaining stiff contact during forward simulation while averaging derivatives from nearby stiff-contact realizations during the backward pass. The method therefore targets gradient variance rather than contact-force modeling error.

3. Local randomized-smoothing mechanism

BCG is activated only when stiff contact is detected. At control step xtx_t2, let xtx_t3 be the set of contacts whose normal impulse exceeds a threshold xtx_t4. A bundle begins when

xtx_t5

The reported experiments use xtx_t6. If no contact exceeds the threshold, the simulator performs an ordinary single-trajectory transition.

For each branch xtx_t7, BCG samples Cartesian perturbations of the contacting chains’ final links:

xtx_t8

Let xtx_t9 denote the Cartesian Jacobian of the contacting links with respect to their joint positions. The perturbations are mapped into joint space using a damped pseudoinverse:

ata_t0

These offsets are embedded into the simulator state,

ata_t1

Only the relevant contacting-chain position and velocity components are changed. The perturbations are treated as constants during differentiation: BCG does not differentiate through random-number generation, the sampled offsets, the Jacobian-based mapping, or the pseudoinverse operation. Consequently,

ata_t2

Each branch is then simulated for ata_t3 control steps using the same finite stiffness ata_t4:

ata_t5

The branch states are aggregated after each bundled transition. For the position and velocity components used in the implementation, aggregation is arithmetic averaging:

ata_t6

At later bundle steps, the policy receives the averaged state,

ata_t7

and the same action is applied to all branches. After ata_t8 steps, ordinary rollout resumes from ata_t9.

Thus, BCG does not average independently optimized trajectories. It constructs a short, coupled bundle in which branches share actions and are repeatedly recombined through the averaged state.

4. Gradient aggregation and optimization

Define the state and action sensitivities by

κ\kappa0

For branch κ\kappa1, let

κ\kappa2

The branch sensitivity recurrence is

κ\kappa3

Because the initial perturbations are constant with respect to policy parameters, all branches begin with the same sensitivity:

κ\kappa4

For arithmetic averaging, the aggregated sensitivity is

κ\kappa5

This averaged sensitivity is the central BCG quantity. It combines transition derivatives generated by nearby stiff-contact configurations before the nominal rollout continues.

If κ\kappa6 denotes a contact-sensitive gradient quantity and κ\kappa7 is a local random perturbation, the construction can be interpreted as the Monte Carlo approximation

κ\kappa8

The method does not define a separate closed-form smoothed objective and does not claim that the estimator is unbiased for the exact nominal nonsmooth contact gradient. The resulting signal is a deliberate bias–variance trade-off: averaging can reduce sensitivity to an individual contact realization, while potentially moving the update away from the exact derivative at the nominal state.

BCG is integrated into a Short Horizon Actor-Critic (SHAC) framework. The actor objective is

κ\kappa9

where ∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.0 is the reward, ∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.1 is a learned critic, ∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.2 is the discount factor, and ∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.3 is the differentiable rollout horizon. The policy update remains

∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.4

BCG changes the state-transition derivative inside this objective; it does not replace SHAC or the actor objective.

5. Relation to contact-gradient and differentiable-contact methods

BCG is related to, but distinct from, several approaches to contact differentiation.

Randomized smoothing for contact optimal control

A prior randomized-smoothing framework defines a bundled gradient as an expectation of local gradients,

∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.5

and applies the same idea to dynamics Jacobians. The corresponding bundled dynamics are

∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.6

with

∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.7

The resulting iterative Randomized Smoothing-MPC method replaces exact iLQR or iMPC linearizations with sampled, smoothed Jacobians (Suh et al., 2021). That framework smooths dynamics around state–input points for trajectory optimization, whereas BCG applies perturbations locally around detected stiff contacts within differentiable policy learning.

Analytic complementarity smoothing

Online friction-coefficient identification has used analytic smoothing of the tangential complementarity condition to obtain nonzero sensitivities near clamping–sliding transitions (Kim et al., 24 Feb 2025). Its smoothed relation is

∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.8

This produces a regularized analytic contact-impulse derivative with respect to friction coefficient. Unlike BCG, it does not average multiple perturbed rollouts; it differentiates a smoothed complementarity relation while retaining hard-contact state prediction.

Overlap-energy and fiber-based gradients

A Fiber Monte Carlo contact method obtains contact forces from an aggregate overlap-energy gradient,

∂xt+1∂θ=∂fκ∂xt∂xt∂θ+∂fκ∂at∂πθ∂θ.\frac{\partial x_{t+1}}{\partial \theta} = \frac{\partial f_\kappa}{\partial x_t} \frac{\partial x_t}{\partial \theta} + \frac{\partial f_\kappa}{\partial a_t} \frac{\partial \pi_\theta}{\partial \theta}.9

where sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},0 is the overlap volume of two bodies (Wang et al., 10 Sep 2025). This is a global geometric gradient rather than a bundle of randomized stiff-contact sensitivities. It provides a BCG-like aggregate contact force conceptually, but it does not implement the BCG policy-learning method.

Differentiable contact-manifold generation

A differentiable contact-manifold framework constructs fixed-size contact bundles using soft top-sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},1 selection, smooth signed distances, witness-point interpolation, and activity weights (Beker et al., 23 Feb 2026). Its fixed-size contacts and activity-weighted aggregation are conceptually compatible with gradient bundling, but the paper does not define or use the BCG acronym. Its primary contribution is smooth, vectorizable contact-manifold generation rather than local randomized smoothing of policy gradients.

6. Experimental evaluation, applications, and limitations

The BCG experiments use a Unitree G1 humanoid, the NVIDIA Warp differentiable rigid-body simulator, Moreau-style time stepping, a modified Gauss–Seidel contact solver, and analytically smoothed contact activation. Four 15-second LAFAN1 motions are studied: Run, Jump, Fight, and Dance. Sim-to-sim evaluation is performed in MuJoCo.

The principal configuration is

sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},2

with perturbation scales

sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},3

and contact threshold sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},4. Branch transitions are executed in parallel on an RTX 2080 Ti GPU. The efficiency discussion reports that BCG experiments can use 64 environments rather than 128 because the gradients are more informative, despite the additional branch computation.

The stiffness comparison uses sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},5. At sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},6, all five trials fall for every motion. At sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},7, no falls are reported for Run, Jump, Fight, or Dance. Mean foot-penetration differences between the training simulator and MuJoCo are also smaller at the higher stiffness. These results support the association between stiff contact, improved contact agreement, and transfer reliability (Aditya et al., 25 Sep 2026).

For the Jump task, the reported gradient analysis finds lower policy-gradient variance with BCG. Individual perturbed branches produce dispersed sensitivity values, while the bundled averages attenuate extreme responses. The supplied results provide directional findings rather than numerical variance values or confidence intervals.

Compared with vanilla SHAC and PPO, SHAC+BCG reaches comparable reward with more than an order of magnitude fewer environment samples than PPO in the reported experiments. Vanilla SHAC stalls at low reward under stiff contact, while BCG achieves higher final reward than PPO for some motions. BCG has greater computational cost than vanilla SHAC because of the branch-based computation graph, although using fewer environments offsets part of this cost.

After transfer to MuJoCo, the reported global mean per-body position errors are:

Motion SHAC-Stiff + BCG PPO
Run sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},8 cm sκ(d)=11+exp⁡(−κd),s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},9 cm
Jump d>0d>00 cm d>0d>01 cm
Fight d>0d>02 cm d>0d>03 cm
Dance d>0d>04 cm d>0d>05 cm

BCG is better on Run, Fight, and Dance by this metric. PPO is slightly better numerically on Jump, although the reported PPO behavior keeps both feet on the ground rather than reproducing the intended jumping behavior.

The learned SHAC+BCG policies are deployed on a real Unitree G1 without hardware fine-tuning. The reported qualitative outcome is successful zero-shot execution of Run, Jump, Fight, and Dance. The supplied material does not specify the hardware control frequency, actuator-level control mode, sensing and state-estimation pipeline, latency compensation, calibration, safety limits, emergency-stop procedure, or sim-to-real parameter randomization.

BCG has several limitations. Its computational and memory overhead grows with the number of branches, especially under sustained contact. Performance depends on the perturbation scales, bundle size, and duration. The method assumes that problematic gradient variation is concentrated near detected contacts and requires a continuous, differentiable finite-stiffness contact model. It is not designed directly for an exactly hard, nonsmooth complementarity simulator. Cartesian perturbations mapped through a damped pseudoinverse may be imperfect near singular configurations or constrained contacts. Arithmetic state averaging may also be inappropriate for all state variables, including orientations, discrete modes, and contact sets. Finally, the averaged derivative is generally not guaranteed to be the exact nominal gradient and may smooth away distinctions between contact modes.

BCG therefore occupies a specific position in differentiable robotics: it preserves stiff forward contact for physical fidelity while replacing a single brittle contact sensitivity with an average over nearby stiff-contact realizations. Its principal contribution is gradient stabilization for first-order policy optimization, not a new contact law, a universal nonsmooth derivative, or a general global smoothing framework.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Bundled Contact Gradients (BCG).