---
title: 'Bundled Contact Gradients (BCG): A Gradient Stabilization Method'
url: https://www.emergentmind.com/topics/bundled-contact-gradients-bcg
type: topic
---

# Bundled Contact Gradients (BCG): A Gradient Stabilization Method

Bundled Contact Gradients (BCG) are a contact-local randomized-smoothing method for differentiable policy learning introduced to stabilize gradients through stiff rigid-body contact. When a stiff contact is detected, BCG evaluates multiple nearby perturbation rollouts around the contacting configuration and averages their state sensitivities before continuing the nominal trajectory. The forward simulator retains finite, stiff contact dynamics, while the gradient signal is locally smoothed to reduce variance caused by contact timing, impact sensitivity, and rapidly changing transition Jacobians [2609.30951]. The acronym BCG is context-dependent: it denotes “Bundled Contact Gradients” in differentiable simulation, but “bundled gradient” has a related meaning in randomized smoothing for contact-aware optimal control [2109.05143], while other literature uses BCG for ballistocardiography and brightest cluster galaxy.

## 1. Terminology and conceptual scope

Bundled Contact Gradients concern the estimation of derivatives through contact-sensitive dynamical systems. They should be distinguished from several unrelated uses of the acronym:

- **Bundled Contact Gradients**: a contact-local randomized-smoothing technique for differentiable policy learning [2609.30951].
- **Bundled gradients in contact optimal control**: smoothed gradients or Jacobians obtained by averaging derivatives over perturbed state–input points, used in iterative randomized-smoothing MPC [2109.05143].
- **Ballistocardiography**: an ambient mechanical biosignal measured from cardiac recoil, used in the BCG-FM foundation model [2606.07692].
- **Brightest cluster galaxy**: the central luminous galaxy in a galaxy cluster, abbreviated BCG in observational astronomy [2303.15557, 2503.21066].

In the differentiable-simulation formulation, the underlying transition is

$$
x_{t+1}=f_\kappa(x_t,a_t),
$$

where $x_t$ is the simulator state, $a_t$ is the action, and $\kappa$ is the contact stiffness. The policy gradient depends on the state sensitivity

$$
\frac{\partial x_{t+1}}{\partial \theta}
=
\frac{\partial f_\kappa}{\partial x_t}
\frac{\partial x_t}{\partial \theta}
+
\frac{\partial f_\kappa}{\partial a_t}
\frac{\partial \pi_\theta}{\partial \theta}.
$$

BCG modifies the derivative evaluation near detected contact events rather than replacing the policy, the forward contact law, or the policy-optimization objective.

The term “bundled” refers to aggregation across a local collection of contact realizations. It does not denote a new physical contact law, a conventional gradient of the nominal trajectory, or a global smoothing operation applied indiscriminately to the simulator.

## 2. Motivation: stiff contact and unstable sensitivities

Differentiable simulation provides analytic derivatives of robot dynamics and can be substantially more sample-efficient than model-free policy-gradient estimators. However, rigid-body contact produces a difficult trade-off between physical fidelity and gradient regularity.

The differentiable simulator considered in the BCG formulation uses an analytically smoothed contact activation,

$$
s_\kappa(d)=\frac{1}{1+\exp(-\kappa d)},
$$

where $d>0$ denotes penetration and $d<0$ denotes separation. Small $\kappa$ produces broad, compliant contact; large finite $\kappa$ produces sharp, stiff contact; and the limit $\kappa\to\infty$ approaches hard-contact-like behavior.

Soft contact is easier to differentiate, but excessive compliance permits unrealistic penetration and delayed or weak force buildup. Policies trained under such dynamics may learn behaviors that do not transfer to hardware. Increasing $\kappa$ improves contact fidelity but makes the transition Jacobians highly sensitive to small perturbations in position and velocity near contact. Nearby contact configurations can then yield sharply different derivatives, resulting in high-variance policy updates, exploding or poorly conditioned sensitivities, dependence on individual contact realizations, and failure to learn dynamic behaviors despite physically appropriate forward dynamics [2609.30951].

The issue is especially consequential for dynamic humanoid motions. Running, jumping, fighting, and dancing depend on foot-ground impulses, impact timing, balance recovery, and whole-body momentum transfer. A compliant model can distort those interactions, whereas a stiff model can make direct backpropagation unreliable.

BCG addresses this trade-off by retaining stiff contact during forward simulation while averaging derivatives from nearby stiff-contact realizations during the backward pass. The method therefore targets gradient variance rather than contact-force modeling error.

## 3. Local randomized-smoothing mechanism

BCG is activated only when stiff contact is detected. At control step $k$, let $\mathcal C_k$ be the set of contacts whose normal impulse exceeds a threshold $\tau$. A bundle begins when

$$
\mathcal C_k\neq\varnothing.
$$

The reported experiments use $\tau=400\,\mathrm{N}$. If no contact exceeds the threshold, the simulator performs an ordinary single-trajectory transition.

For each branch $i\in\{1,\ldots,B\}$, BCG samples Cartesian perturbations of the contacting chains’ final links:

$$
\Delta p^{(i)}\sim\mathcal N(0,\sigma_p^2I),
\qquad
\Delta v^{(i)}\sim\mathcal N(0,\sigma_v^2I).
$$

Let $J_k$ denote the Cartesian Jacobian of the contacting links with respect to their joint positions. The perturbations are mapped into joint space using a damped pseudoinverse:

$$
\Delta q^{(i)}=J_k^\dagger\Delta p^{(i)},
\qquad
\Delta\dot q^{(i)}=J_k^\dagger\Delta v^{(i)}.
$$

These offsets are embedded into the simulator state,

$$
x_k^{(i)}=x_k+\delta x_k^{(i)}.
$$

Only the relevant contacting-chain position and velocity components are changed. The perturbations are treated as constants during differentiation: BCG does not differentiate through random-number generation, the sampled offsets, the Jacobian-based mapping, or the pseudoinverse operation. Consequently,

$$
\frac{d x_k^{(i)}}{d\theta}
=
\frac{d x_k}{d\theta}.
$$

Each branch is then simulated for $H$ control steps using the same finite stiffness $\kappa$:

$$
x_{t+1}^{(i)}
=
f_\kappa(x_t^{(i)},a_t),
\qquad k\leq t<k+H.
$$

The branch states are aggregated after each bundled transition. For the position and velocity components used in the implementation, aggregation is arithmetic averaging:

$$
\bar{x}_{t+1}
=
\frac{1}{B}
\sum_{i=1}^{B}x_{t+1}^{(i)}.
$$

At later bundle steps, the policy receives the averaged state,

$$
a_t=\pi_\theta(\bar{x}_t),
\qquad t>k,
$$

and the same action is applied to all branches. After $H$ steps, ordinary rollout resumes from $\bar{x}_{k+H}$.

Thus, BCG does not average independently optimized trajectories. It constructs a short, coupled bundle in which branches share actions and are repeatedly recombined through the averaged state.

## 4. Gradient aggregation and optimization

Define the state and action sensitivities by

$$
Z_t=\frac{d x_t}{d\theta},
\qquad
U_t=\frac{d a_t}{d\theta}.
$$

For branch $i$, let

$$
F_{x,t}^{(i)}
=
\left.
\frac{\partial f_\kappa}{\partial x}
\right|_{(x_t^{(i)},a_t)},
\qquad
F_{a,t}^{(i)}
=
\left.
\frac{\partial f_\kappa}{\partial a}
\right|_{(x_t^{(i)},a_t)}.
$$

The branch sensitivity recurrence is

$$
Z_{t+1}^{(i)}
=
F_{x,t}^{(i)}Z_t^{(i)}
+
F_{a,t}^{(i)}U_t.
$$

Because the initial perturbations are constant with respect to policy parameters, all branches begin with the same sensitivity:

$$
Z_k^{(i)}=Z_k.
$$

For arithmetic averaging, the aggregated sensitivity is

$$
\bar Z_{t+1}
=
\frac{1}{B}
\sum_{i=1}^{B}
\left(
F_{x,t}^{(i)}Z_t^{(i)}
+
F_{a,t}^{(i)}U_t
\right).
$$

This averaged sensitivity is the central BCG quantity. It combines transition derivatives generated by nearby stiff-contact configurations before the nominal rollout continues.

If $g(x)$ denotes a contact-sensitive gradient quantity and $\delta x$ is a local random perturbation, the construction can be interpreted as the Monte Carlo approximation

$$
\mathbb E_{\delta x}[g(x+\delta x)]
\approx
\frac{1}{B}
\sum_{i=1}^{B}g(x+\delta x^{(i)}).
$$

The method does not define a separate closed-form smoothed objective and does not claim that the estimator is unbiased for the exact nominal nonsmooth contact gradient. The resulting signal is a deliberate bias–variance trade-off: averaging can reduce sensitivity to an individual contact realization, while potentially moving the update away from the exact derivative at the nominal state.

BCG is integrated into a Short Horizon Actor-Critic (SHAC) framework. The actor objective is

$$
\mathcal L_{\mathrm{FO}}(\theta)
=
-\sum_{t=t_0}^{t_0+N-1}
\gamma^{t-t_0}r(x_t,a_t)
-
\gamma^N V_\psi(x_{t_0+N}),
$$

where $r$ is the reward, $V_\psi$ is a learned critic, $\gamma$ is the discount factor, and $N$ is the differentiable rollout horizon. The policy update remains

$$
\theta
\leftarrow
\theta-\eta_\theta
\nabla_\theta\mathcal L_{\mathrm{FO}}.
$$

BCG changes the state-transition derivative inside this objective; it does not replace SHAC or the actor objective.

## 5. Relation to contact-gradient and differentiable-contact methods

BCG is related to, but distinct from, several approaches to contact differentiation.

### Randomized smoothing for contact optimal control

A prior randomized-smoothing framework defines a bundled gradient as an expectation of local gradients,

$$
\widetilde{\nabla}f(x)
=
\mathbb E_{w\sim p}
[\nabla f(x+w)],
$$

and applies the same idea to dynamics Jacobians. The corresponding bundled dynamics are

$$
\widetilde f(x,u)
=
\mathbb E_{w,v}[f(x+w,u+v)],
$$

with

$$
\widetilde A
=
\frac{\partial\widetilde f}{\partial x},
\qquad
\widetilde B
=
\frac{\partial\widetilde f}{\partial u}.
$$

The resulting iterative Randomized Smoothing-MPC method replaces exact iLQR or iMPC linearizations with sampled, smoothed Jacobians [2109.05143]. That framework smooths dynamics around state–input points for trajectory optimization, whereas BCG applies perturbations locally around detected stiff contacts within differentiable policy learning.

### Analytic complementarity smoothing

Online friction-coefficient identification has used analytic smoothing of the tangential complementarity condition to obtain nonzero sensitivities near clamping–sliding transitions [2502.16843]. Its smoothed relation is

$$
\|v_k^t\|
\left(
\mu^2(\lambda_k^n)^2
-
\|\lambda_k^t\|^2
\right)
=
\rho_t.
$$

This produces a regularized analytic contact-impulse derivative with respect to friction coefficient. Unlike BCG, it does not average multiple perturbed rollouts; it differentiates a smoothed complementarity relation while retaining hard-contact state prediction.

### Overlap-energy and fiber-based gradients

A Fiber Monte Carlo contact method obtains contact forces from an aggregate overlap-energy gradient,

$$
\boldsymbol F_c
=
-\nabla_{\boldsymbol u}w(V_c),
$$

where $V_c$ is the overlap volume of two bodies [2509.08609]. This is a global geometric gradient rather than a bundle of randomized stiff-contact sensitivities. It provides a BCG-like aggregate contact force conceptually, but it does not implement the BCG policy-learning method.

### Differentiable contact-manifold generation

A differentiable contact-manifold framework constructs fixed-size contact bundles using soft top-$K$ selection, smooth signed distances, witness-point interpolation, and activity weights [2602.20304]. Its fixed-size contacts and activity-weighted aggregation are conceptually compatible with gradient bundling, but the paper does not define or use the BCG acronym. Its primary contribution is smooth, vectorizable contact-manifold generation rather than local randomized smoothing of policy gradients.

## 6. Experimental evaluation, applications, and limitations

The BCG experiments use a Unitree G1 humanoid, the NVIDIA Warp differentiable rigid-body simulator, Moreau-style time stepping, a modified Gauss–Seidel contact solver, and analytically smoothed contact activation. Four 15-second LAFAN1 motions are studied: Run, Jump, Fight, and Dance. Sim-to-sim evaluation is performed in MuJoCo.

The principal configuration is

$$
N=32,\qquad B=10,\qquad H=2,
$$

with perturbation scales

$$
\sigma_p=1\,\mathrm{cm},
\qquad
\sigma_v=2\,\mathrm{cm/s},
$$

and contact threshold $\tau=400\,\mathrm{N}$. Branch transitions are executed in parallel on an RTX 2080 Ti GPU. The efficiency discussion reports that BCG experiments can use 64 environments rather than 128 because the gradients are more informative, despite the additional branch computation.

The stiffness comparison uses $\kappa=50,100,300$. At $\kappa=50$, all five trials fall for every motion. At $\kappa=300$, no falls are reported for Run, Jump, Fight, or Dance. Mean foot-penetration differences between the training simulator and MuJoCo are also smaller at the higher stiffness. These results support the association between stiff contact, improved contact agreement, and transfer reliability [2609.30951].

For the Jump task, the reported gradient analysis finds lower policy-gradient variance with BCG. Individual perturbed branches produce dispersed sensitivity values, while the bundled averages attenuate extreme responses. The supplied results provide directional findings rather than numerical variance values or confidence intervals.

Compared with vanilla SHAC and PPO, SHAC+BCG reaches comparable reward with more than an order of magnitude fewer environment samples than PPO in the reported experiments. Vanilla SHAC stalls at low reward under stiff contact, while BCG achieves higher final reward than PPO for some motions. BCG has greater computational cost than vanilla SHAC because of the branch-based computation graph, although using fewer environments offsets part of this cost.

After transfer to MuJoCo, the reported global mean per-body position errors are:

| Motion | SHAC-Stiff + BCG | PPO |
|---|---:|---:|
| Run | $34.1\pm0.4$ cm | $42.8\pm0.7$ cm |
| Jump | $18.6\pm0.4$ cm | $18.0\pm0.2$ cm |
| Fight | $13.5\pm0.3$ cm | $34.3\pm1.4$ cm |
| Dance | $18.5\pm0.8$ cm | $19.0\pm0.1$ cm |

BCG is better on Run, Fight, and Dance by this metric. PPO is slightly better numerically on Jump, although the reported PPO behavior keeps both feet on the ground rather than reproducing the intended jumping behavior.

The learned SHAC+BCG policies are deployed on a real Unitree G1 without hardware fine-tuning. The reported qualitative outcome is successful zero-shot execution of Run, Jump, Fight, and Dance. The supplied material does not specify the hardware control frequency, actuator-level control mode, sensing and state-estimation pipeline, latency compensation, calibration, safety limits, emergency-stop procedure, or sim-to-real parameter randomization.

BCG has several limitations. Its computational and memory overhead grows with the number of branches, especially under sustained contact. Performance depends on the perturbation scales, bundle size, and duration. The method assumes that problematic gradient variation is concentrated near detected contacts and requires a continuous, differentiable finite-stiffness contact model. It is not designed directly for an exactly hard, nonsmooth complementarity simulator. Cartesian perturbations mapped through a damped pseudoinverse may be imperfect near singular configurations or constrained contacts. Arithmetic state averaging may also be inappropriate for all state variables, including orientations, discrete modes, and contact sets. Finally, the averaged derivative is generally not guaranteed to be the exact nominal gradient and may smooth away distinctions between contact modes.

BCG therefore occupies a specific position in differentiable robotics: it preserves stiff forward contact for physical fidelity while replacing a single brittle contact sensitivity with an average over nearby stiff-contact realizations. Its principal contribution is gradient stabilization for first-order policy optimization, not a new contact law, a universal nonsmooth derivative, or a general global smoothing framework.

Source: https://www.emergentmind.com/topics/bundled-contact-gradients-bcg