Bundled Contact Gradients: Stabilizing Differentiable Simulation for Deployable Dynamic Tasks
Abstract: Differentiable simulation provides analytic gradients of robot dynamics, enabling fast and sample-efficient first-order policy optimization. However, obtaining smooth and informative gradients through rigid-body contact typically requires softened contact models, often at the expense of physical fidelity and thereby limiting learned policies largely to simulation. This trade-off becomes particularly consequential for dynamic humanoid motions, where accurate contact dynamics are critical for transferring policies to the real world. Increasing contact stiffness in rigid-body simulation improves the fidelity of interactions, but also makes the dynamics increasingly sensitive to small state perturbations, producing high-variance gradients that can destabilize first-order policy learning. To address this, we propose \emph{Bundled Contact Gradients (BCG)}, a contact-local randomized smoothing framework for differentiable policy learning. When stiff contact is detected, our method evaluates a local bundle of randomized perturbation rollouts around the stiff contact configuration and aggregates their gradient signal thereby reducing gradient variance. We demonstrate the effectiveness of our method by successfully training and transferring dynamic motions zero-shot onto a real-world Unitree G1 humanoid platform. Videos and supplementary information can be found at https://bundledcontactgradients.github.io/
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper presents a new method for teaching humanoid robots to perform fast, complicated movements, such as running, jumping, fighting, and dancing.
The researchers want robots to learn these movements in a computer simulation and then perform them on a real robot. The main challenge is that robots often touch the ground. These contacts must be modeled accurately, but accurate contact simulations can make learning unstable.
The paper introduces a method called Bundled Contact Gradients, or BCG, to solve this problem.
2. What questions are the researchers asking?
The researchers are mainly asking:
- Can a robot learn realistic movements while using accurate, stiff ground contact in simulation?
- Can BCG make the learning process more stable?
- Does BCG require fewer examples and less training time than common reinforcement-learning methods?
- Do the learned movements still work in another simulator?
- Can the movements transfer directly to a real humanoid robot without extra training?
The robot used in the experiments is the Unitree G1, a human-shaped robot.
3. How did the researchers approach the problem?
Teaching robots through trial and error
The researchers use reinforcement learning, which is similar to training a dog with rewards. The robot tries an action, such as moving a leg, and receives a better reward when its movement is closer to the target motion.
The target motions came from a motion-capture dataset containing examples of running, jumping, fighting, and dancing.
Using a differentiable simulator
The robot first learns inside a physics simulator. A simulator is a computer program that predicts how the robot will move.
This simulator is differentiable, meaning it can calculate how a small change in an action might change the final result. For example, it can estimate:
“If the robot moves its foot slightly faster, how will that affect its balance later?”
These calculations are called gradients. They help the learning system choose better actions quickly, much like using a map showing which direction leads uphill.
This approach can be much more efficient than trying completely random actions. However, it becomes difficult when the robot hits the ground.
Why ground contact causes trouble
When a foot touches the ground, a tiny change in position or speed can cause a large change in the robot’s motion. For example, a foot might:
- bounce,
- slide,
- push strongly against the ground, or
- miss the ground completely.
With very stiff, realistic contact, the calculated gradients can point in very different directions. This confuses the learning system and can cause training to fail.
Using softer contact makes learning easier, but the robot may learn unrealistic behavior, such as its feet sinking too far into the floor. Such a policy may work in simulation but fail on a real robot.
The BCG idea
BCG deals with this problem by creating a small group, or bundle, of nearby possibilities whenever a strong contact is detected.
For example, if a foot hits the ground, the simulator creates several slightly different versions of that event:
- one where the foot is a little higher,
- one where it is a little lower,
- one where it is moving slightly faster, and so on.
The simulator runs all these versions at the same time. It then averages their results and their gradients.
An everyday analogy is trying to decide which way to steer a bicycle on a rough path. Instead of trusting one possibly misleading measurement, you test several nearby paths and use their average direction. This produces a steadier decision.
Importantly, BCG uses accurate, stiff contact during the forward simulation. It does not simply make the ground soft. It only smooths the information used for learning.
The researchers combine BCG with:
- SHAC, a method that learns using gradients from short simulation periods; and
- ADD, a system that compares the robot’s movement with the target motion and produces a reward.
4. What did the researchers find?
The researchers tested four motions: Run, Jump, Fight, and Dance.
Stiff contact improved realism and transfer
When the simulator used soft contact, the robot often failed after being transferred to another simulator. With stiffer contact, the learned movements were more realistic and reliable.
At the highest tested stiffness, the robot had no falls in the reported tests. Softer settings caused many falls, especially after transfer to MuJoCo, another physics simulator.
BCG reduced unstable gradients
Without BCG, small differences in contact situations produced very different gradients. This made the learning updates inconsistent.
With BCG, the gradients were averaged across several nearby contact situations. The gradients became more similar and stable. This gave the learning system a clearer signal about how to improve the robot’s movements.
BCG learned more efficiently than PPO
The researchers compared BCG with PPO, a widely used reinforcement-learning method.
BCG achieved similar reward levels using more than ten times fewer simulated experiences in some experiments. In other words, it learned more from each example.
The regular differentiable method without BCG, called vanilla SHAC, often stopped at a low reward. This suggests that accurate contact alone was not enough; the gradients also needed to be stabilized.
BCG required somewhat more computation for each simulation because it runs several nearby versions of contact. However, it could use fewer total environments and still learn effectively.
BCG improved motion tracking
After the policies were moved from the training simulator to MuJoCo, BCG had lower average tracking error than PPO on three of the four motions.
| Motion | BCG tracking error | PPO tracking error |
|---|---|---|
| Run | 34.1 cm | 42.8 cm |
| Jump | 18.6 cm | 18.0 cm |
| Fight | 13.5 cm | 34.3 cm |
| Dance | 18.5 cm | 19.0 cm |
Although PPO had a slightly smaller numerical error for Jump, the paper notes that PPO did not really reproduce the jump. Instead, it kept both feet on the ground. This shows that a simple error number does not always tell the whole story.
The movements worked on a real robot
Finally, the researchers placed the learned policies directly onto a real Unitree G1 robot. The robot successfully performed the learned dynamic motions without extra training or hardware-specific adjustment.
This is called zero-shot sim-to-real transfer: the policy goes from simulation to the real world in one step.
5. Why are these results important?
The paper shows that differentiable simulation can be useful even for difficult robot movements involving strong impacts and ground contact.
Previously, researchers often had to choose between:
- realistic contact, which makes learning unstable; or
- smooth contact, which makes learning easier but can produce unrealistic behaviors.
BCG offers a compromise. It keeps the simulation physically realistic while making the learning signal more reliable.
This could make it easier to train robots using fewer computer simulations and less time. It may also help robots learn other difficult skills, such as climbing, lifting objects, manipulating tools, or performing acrobatic movements.
However, BCG is not perfect. Running several contact variations requires extra memory and computation. The researchers must also choose suitable settings, such as how many variations to create and how large the small movements should be.
Simple conclusion
The main idea of the paper is:
When a robot’s foot hits the ground, do not trust just one exact contact event. Try several nearby versions, average the information, and use that steadier signal to teach the robot.
Using this method, the researchers trained a humanoid robot to learn dynamic movements more efficiently and perform them on real hardware. The work suggests that better handling of contact could help bridge the gap between robot learning in simulation and reliable behavior in the real world.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited task diversity: The evaluation covers only four prerecorded LAFAN1 motions—Run, Jump, Fight, and Dance—so it remains unclear whether BCG generalizes to locomotion, recovery from disturbances, loco-manipulation, dexterous manipulation, or task-driven behaviors.
- Single robot platform: All experiments use the Unitree G1, leaving the method’s applicability to different humanoid morphologies, actuator configurations, mass distributions, and contact geometries unresolved.
- Narrow hardware evidence: Zero-shot sim-to-real transfer is presented primarily as qualitative video evidence. The paper does not report systematic hardware metrics such as tracking error, fall rate over many trials, energy consumption, actuator saturation, latency sensitivity, or robustness to disturbances.
- Insufficient statistical validation: Most evaluations use only five runs, and the hardware experiments do not appear to include repeated trials or confidence intervals. The reliability and statistical significance of the reported improvements therefore remain uncertain.
- No controlled comparison with alternative smoothing methods: BCG is not directly compared against global randomized smoothing, analytically softened contacts, adaptive-horizon methods such as AHAC, implicit or complementarity-based differentiation, or other gradient-variance-reduction techniques.
- Incomplete ablation of BCG components: The experiments do not isolate the effects of bundle size , bundle duration , position perturbations, velocity perturbations, contact-triggered activation, or state aggregation. It is therefore unclear which components are necessary for the observed gains.
- Unresolved hyperparameter sensitivity: The perturbation scales and , contact threshold , stiffness , bundle size, and bundle duration are manually selected. The paper does not characterize performance degradation when these values are misspecified.
- No principled rule for perturbation scale selection: The relationship between perturbation magnitude, local contact geometry, simulator timestep, robot velocity, and the bias introduced by smoothing is not derived or empirically mapped.
- Potential smoothing bias is not quantified: Averaging branch gradients produces a gradient for a locally smoothed system, not necessarily the original stiff-contact objective. The paper does not measure how closely the BCG gradient aligns with the unsmoothed gradient or how smoothing changes the learned policy’s objective.
- Forward-dynamics inconsistency is insufficiently analyzed: Although the paper states that forward contact remains stiff, branch states are averaged before subsequent policy evaluation and rollout continuation. The physical interpretation and stability of this averaged state—especially for orientations, angular velocities, contact modes, and constraints—are not established.
- Aggregation is limited to arithmetic means: The method averages positions and velocities, but the paper does not explain how orientations, contact states, impulses, or other non-Euclidean and constraint-dependent quantities are handled in general.
- Perturbations may violate physical constraints: Cartesian perturbations mapped through a damped pseudoinverse can produce joint configurations, velocities, self-collisions, or contact states that are physically implausible. The frequency and effect of such invalid branches are not reported.
- Dependence on Jacobian conditioning is unexplored: The stability of the damped pseudoinverse near singular configurations, during multi-contact interactions, or for coupled contacting chains is not analyzed.
- Multi-contact and contact-mode scalability are unclear: The method is described for the set of contacts exceeding a force threshold, but the behavior of perturbation sampling and aggregation under simultaneous feet, hands, knees, or object contacts remains untested.
- Contact detection introduces a discontinuity: BCG activation depends on a hard impulse threshold . The resulting switching behavior, sensitivity to small force fluctuations, and effects on training stability are not investigated.
- Sustained contact overhead is not measured: The conclusion acknowledges increased computation and memory for sustained stiff contact, but the paper does not provide detailed runtime, memory, GPU utilization, or scaling results as contact frequency and bundle size increase.
- Computational comparisons are not normalized: Claims of efficiency do not fully separate environment samples, simulator evaluations, differentiable branch evaluations, GPU time, memory usage, and wall-clock cost. Comparisons with PPO may therefore use different effective computational budgets.
- No scaling study with bundle size: The reported configuration uses , but the trade-off between gradient variance, bias, memory, and runtime for different values of is not shown.
- No scaling study with bundle horizon: The use of leaves open whether longer bundled horizons improve contact sensitivity estimation or instead amplify gradient instability and computational cost.
- Gradient-quality evaluation is indirect: Gradient variance is measured by summing component-wise sample variances across environments, but the paper does not evaluate gradient bias, cosine alignment with finite-difference estimates, update effectiveness, or variance relative to PPO and other baselines under matched conditions.
- The variance metric may conflate environment heterogeneity with estimator noise: The reported reduction in variance across environments does not distinguish stochastic variation caused by randomized bundles from differences in states, contacts, motion phases, or task difficulty.
- Randomness and reproducibility are underspecified: The paper does not report the number of independent training seeds, randomization distributions in full detail, or whether the same sampled perturbations are used across compared methods.
- Limited stiffness range: The experiments evaluate only and do not establish behavior near the hard-contact limit or across a broader range of simulator timesteps and stiffness values.
- Physical fidelity is measured narrowly: Sim-to-sim transfer is assessed mainly through falls and foot penetration depth. Other contact quantities—impulses, frictional forces, slip, impact timing, joint loads, and center-of-mass dynamics—are not compared between simulators or against hardware.
- MuJoCo transfer does not establish real-world model validity: Agreement with MuJoCo may reflect compatibility between two simulators rather than accuracy with respect to real contact dynamics. Direct validation against measured hardware contacts is absent.
- Domain randomization is not characterized: The paper does not specify whether dynamics, friction, actuator properties, latency, sensing noise, terrain, or state-estimation errors are randomized during training, making the source of sim-to-real robustness unclear.
- Robustness to disturbances is untested: The learned policies are not evaluated under pushes, uneven terrain, changes in friction, external contacts, payload changes, or actuator degradation.
- The role of ADD is confounded with BCG: Because BCG is evaluated primarily within an ADD-based imitation framework, it is unclear whether its benefits persist with conventional tracking rewards, task rewards, sparse rewards, or other discriminators.
- Reward and critic interactions are underexplored: The effect of BCG on critic accuracy, bootstrapping error, reward-gradient scale, and actor–critic instability is not reported.
- Long-horizon behavior remains unresolved: SHAC still uses a short differentiable horizon of and a critic for the remaining return. The method’s effectiveness for longer-horizon dependencies and delayed contact consequences is unknown.
- Failure cases are not reported: The paper does not analyze motions or training runs in which BCG fails, diverges, produces unstable policies, or converges to physically plausible but behaviorally incorrect solutions.
- The claimed causal mechanism is not fully established: The results show lower measured gradient variability and improved task outcomes, but they do not demonstrate that the improvement is specifically caused by better contact gradients rather than altered exploration, state averaging, regularization, or effective batch-size changes.
- Action-sharing across branches may limit validity: All perturbed branches use shared actions, even after their states diverge. The consequences for feedback control, branch-wise action adaptation, and the accuracy of the resulting policy gradient are not examined.
- Perturbations are treated as nondifferentiable constants: Ignoring derivatives through the Jacobian-based perturbation map and sampling process simplifies implementation but may omit useful sensitivities and makes the resulting estimator’s theoretical interpretation incomplete.
- No convergence or optimization theory is provided: The conditions under which local bundle averaging reduces variance without preventing convergence, and how the estimator relates to the gradient of a smoothed objective, remain formally unresolved.
- Contact-model dependence is unknown: BCG is built on one analytically smoothed Moreau-style contact model. Its effectiveness with penalty contacts, compliant contact, complementarity solvers, implicit contact models, or other differentiable simulators has not been demonstrated.
- Hardware safety and deployment constraints are insufficiently documented: The paper does not detail safeguards, torque limits, emergency stopping criteria, controller latency, state-estimation architecture, or the extent to which these factors constrain successful real-world deployment.
- Generalization beyond motion imitation is unverified: The conclusion proposes higher-dimensional loco-manipulation and dexterous manipulation, but no evidence establishes that contact-local perturbations remain effective when contacts involve objects, frictional manipulation, grasp transitions, or highly discontinuous contact modes.
Practical Applications
Immediate Applications
- Deployable dynamic control for humanoid robots — robotics and manufacturing.
BCG can be integrated into GPU-based differentiable simulation pipelines to train policies for running, jumping, dancing, whole-body motion, and rapid balance recovery while retaining relatively stiff, physically realistic foot–ground contact. The demonstrated workflow is: reference motion dataset →
ADDimitation reward →SHAC + BCGpolicy optimization → sim-to-sim validation → hardware deployment. The paper reports successful zero-shot execution on a Unitree G1, making this relevant to research platforms, warehouse humanoids, inspection robots, and entertainment or demonstration robots. Dependencies: a sufficiently accurate robot model, actuator and friction calibration, safety constraints, reliable state estimation, and a simulator supporting differentiable contact dynamics. The reported hardware evidence is limited to one humanoid platform and four motions. - More sample-efficient training for contact-rich robot skills — industrial robotics.
Robot developers can use BCG as an alternative or supplement to PPO when training involves repeated impacts, foot contacts, or fast changes in support. The experiments indicate that
SHAC + BCGcan achieve comparable reward with more than an order of magnitude fewer environment samples than PPO, although with additional per-sample computation. This may reduce simulation time and the number of GPU-hours required for early-stage controller development. Dependencies: the task must be expressible through differentiable or approximately differentiable dynamics; GPU parallelism is important because contact branches are simulated concurrently. The computational benefit depends on how often stiff-contact bundling is activated. - Motion imitation from human demonstrations — animation, entertainment, and rehabilitation robotics. LAFAN1-style motion data can be converted into dynamic robot behaviors using an adversarial differential discriminator. Potential products include tools that retarget human motion to humanoid robots, generate physically executable animation previews, or train robotic exercise and rehabilitation assistants to reproduce therapist-designed movements. BCG is especially useful when the motion includes jumps, impacts, or rapid foot placement that soft-contact simulators cannot reproduce faithfully. Dependencies: reference motions must be compatible with the robot’s morphology and joint limits; the discriminator and reward design must avoid rewarding visually similar but unsafe behavior; hardware validation remains necessary.
- A reusable simulator component for stiff-contact policy learning — software and simulation infrastructure.
BCG can be implemented as a contact-triggered module in frameworks such as NVIDIA Warp or similar differentiable physics systems. A practical implementation would expose parameters such as bundle size
B, bundle durationH, contact thresholdτ, and Cartesian perturbation scalesσpandσv, allowing users to trade off gradient stability, bias, memory, and runtime. Dependencies: the contact model must remain continuous with computable local derivatives. BCG is not directly applicable to truly discontinuous contact transitions without additional smoothing or specialized differentiation. - Contact-gradient diagnostics for robotics research and debugging — academia and R&D. Researchers can use the branch-gradient spread as a diagnostic for identifying unstable contact configurations, problematic simulator stiffness, or conflicting policy-update directions. Comparing individual branch sensitivities with their bundled average can reveal whether training failures arise from contact sensitivity rather than from the policy architecture or reward. This supports reproducible ablation studies and principled tuning of contact parameters. Dependencies: gradient statistics must be logged consistently, and reduced variance should not be mistaken for improved physical correctness; smoothing can introduce bias.
- Faster sim-to-sim validation workflows — robotics engineering. Policies trained with stiff contact can be tested in a second physics engine, such as MuJoCo, before hardware trials. This provides an intermediate validation stage for foot penetration, falls, tracking error, and sensitivity to contact-model differences. The paper reports lower tracking error than PPO on three of four motions after transfer to MuJoCo. Dependencies: agreement between simulators is not sufficient evidence of hardware safety. Validation should include randomized masses, friction, delays, sensor noise, actuator limits, and terrain conditions.
- Educational and laboratory platforms for differentiable robotics — academia. The method provides a concrete teaching and research example connecting automatic differentiation, randomized smoothing, reinforcement learning, contact mechanics, and sim-to-real transfer. A laboratory workflow could compare vanilla SHAC, SHAC with soft contact, SHAC with stiff contact, BCG, and PPO on the same robot model. Dependencies: access to GPU simulation, differentiable physics software, and suitable hardware or benchmark environments. Results may be sensitive to implementation details and hyperparameter tuning.
- Safety-oriented offline testing of dynamic behaviors — policy and robotics operations. Organizations deploying legged or humanoid robots can use BCG-trained controllers in a digital testbed to evaluate fall likelihood, contact-force excursions, foot penetration, and robustness before approving physical tests. The paper’s contact-stiffness ablation illustrates how simulation contact fidelity affects transfer reliability. Dependencies: the testbed must include conservative uncertainty models and independently defined safety thresholds. BCG itself does not provide formal safety guarantees or replace runtime monitoring.
Long-Term Applications
- Whole-body loco-manipulation — logistics, construction, and service robotics. Extending BCG from foot–ground contacts to simultaneous hand, foot, object, and environment contacts could enable humanoids to carry loads, open doors, climb, push objects, or recover from disturbances while maintaining dynamic motion. A future controller could trigger separate local bundles for each active contact chain and coordinate them through shared policy actions. Dependencies: scalable handling of multiple simultaneous contacts, accurate friction and object models, collision detection, and computational methods that prevent branch-count growth from becoming prohibitive.
- Dexterous manipulation and high-impact grasping — robotics and prosthetics. Localized gradient averaging could stabilize learning for grasp transitions, in-hand manipulation, tool use, throwing, catching, and contact-rich assembly. The Cartesian perturbation formulation is potentially compatible with fingertips, palms, and tool contact points rather than only robot feet. Dependencies: contact geometry is higher-dimensional and often includes frictional transitions, rolling, slip, and deformable objects. The current experiments do not establish effectiveness for manipulation or dexterous hands.
- Adaptive, sensitivity-aware bundle selection — autonomous robot software.
Future systems could automatically vary
B,H,σp, andσvbased on local gradient disagreement, contact force, Jacobian conditioning, or estimated model uncertainty. Stable contacts could use few or no branches, whereas highly sensitive impacts could receive larger bundles. This would reduce overhead while preserving smoothing where it is most useful. Dependencies: reliable online sensitivity metrics, bounded computational latency, and safeguards against excessive smoothing or unstable adaptive feedback. - Robust policy training across hardware uncertainty — industrial deployment. BCG could be combined with domain randomization and system identification to train policies robust to variation in friction, payload, actuator strength, terrain compliance, sensor latency, and joint calibration. A productized workflow might jointly optimize the controller and uncertain simulator parameters using contact-local gradients. Dependencies: randomized contact gradients may average away real but important failure modes; uncertainty distributions must be measured from hardware rather than chosen arbitrarily. Hardware-in-the-loop validation remains necessary.
- Dynamic control of quadrupeds, aerial robots, and multi-physics systems — broader autonomy. The underlying idea may transfer to quadruped jumping, agile flight with intermittent impacts, wheeled–legged vehicles, and systems involving fluid, elastic, or compliant dynamics. In these settings, contact-local smoothing could complement differentiable multiphysics simulators and reduce instability in first-order policy optimization. Dependencies: each domain has different nonsmooth phenomena, such as aerodynamic stall, wheel slip, or deformation. The paper’s results cannot be assumed to generalize without domain-specific contact detection, perturbation mappings, and validation.
- Fast model-based design optimization — robot and mechanism engineering. Because BCG retains gradients through relatively stiff contact, it could support optimization of foot geometry, compliance, mass distribution, linkage dimensions, gait timing, or actuator placement alongside policy parameters. This could produce robot designs that are easier to control dynamically and transfer more reliably to hardware. Dependencies: differentiating through geometry, collisions, and changing contact topology is more difficult than differentiating policy parameters. Design gradients may also be biased by the randomized smoothing procedure.
- Closed-loop digital twins for fleet management — industry and infrastructure. Once validated, a differentiable contact model could be used in a digital twin to simulate robot-specific wear, payload changes, terrain conditions, and controller updates. BCG could help periodically retrain or adapt policies without requiring large quantities of physical trial data. Dependencies: this requires high-fidelity calibration from operational data, secure data pipelines, validated uncertainty estimates, and strict separation between simulation recommendations and safety-critical execution.
- Human-assistive and rehabilitation systems — healthcare. Dynamic motion imitation could eventually support powered exoskeletons, lower-limb rehabilitation devices, and humanoid assistants that reproduce therapist-prescribed movements while adapting to patient-specific contact and balance conditions. BCG may be useful for learning stable transitions involving foot placement and body support. Dependencies: medical certification, interpretable safety limits, patient-specific biomechanics, fail-safe control, and extensive clinical validation are required. The current study provides no clinical evidence.
- Robotics policy benchmarking and standardized tooling — academia and policy research. The method could motivate benchmarks that report not only reward, but also gradient variance, contact stiffness, sample efficiency, sim-to-sim error, hardware transfer rate, and computational cost. Such benchmarks would help institutions and regulators compare claims about “sim-to-real” learning more transparently. Dependencies: standardized robot models, reference motions, hardware protocols, and independently reproducible implementations are needed. Single-platform demonstrations are insufficient for broad policy conclusions.
- Real-time adaptation and recovery from unexpected contact — service and field robotics. A future deployment system might use local contact perturbations during online model-predictive control or policy refinement to handle slips, uneven terrain, collisions, and unmodeled obstacles. The same variance-reduction principle could help estimate useful local sensitivities during rapid recovery maneuvers. Dependencies: current BCG is presented primarily as an offline training method. Real-time use would require strict latency bounds, bounded memory, online-safe updates, and guarantees that perturbation averaging does not obscure rare but critical failure states.
Glossary
- Adversarial Differential Discriminators (ADD): Learned discriminators that generate differentiable rewards for motion imitation by distinguishing reference or ideal motion features from policy-generated ones. “we apply this framework to motion imitation, where the policy tracks reference motions using rewards from Adversarial Differential Discriminators (ADD)”
- Analytic gradient: A gradient computed from explicit derivatives of a differentiable model rather than estimated from sampled perturbations. “This first-order gradient is typically far lower in variance than the score-function estimator”
- Automatic differentiation: A computational method that evaluates derivatives by systematically applying the chain rule through a program. “the backward pass with automatic differentiation”
- Backpropagation: Reverse-mode differentiation through a sequence of computations, used here to propagate policy gradients through simulated dynamics. “by backpropagating gradients through the simulated dynamics to the policy”
- Bundle gradient: A gradient obtained by combining derivatives from multiple nearby perturbed trajectories or contact configurations. “Bundled gradients replace the exact contact derivative with a randomized-smoothing estimate”
- Cartesian space: A coordinate representation describing the position and motion of objects in physical three-dimensional space. “the simulator evaluates a small bundle of perturbations sampled in Cartesian space”
- Contact-implicit trajectory optimization: Trajectory optimization that incorporates contact events and constraints directly into the optimization problem rather than prescribing them in advance. “complementing classical contact-implicit trajectory optimization”
- Contact impulse: A short-duration change in momentum caused by a collision or contact interaction. “A bundle step is triggered when the set of contacts whose normal impulse exceeds a detection threshold is nonempty.”
- Contact stiffness: A parameter describing how strongly a contact model responds to penetration or deformation. “The parameter controls the extent of smoothing (i.e. the contact stiffness)”
- Complementarity constraint: A mathematical constraint requiring two quantities, such as contact force and separation, to satisfy mutually exclusive conditions. “Hard-contact models instead express non-penetration and friction through complementarity constraints”
- Differentiable rollout: A simulated sequence of states and actions through which derivatives can be computed. “SHAC+BCG experiments use a differentiable rollout horizon of control steps”
- Damped pseudoinverse: A numerically stabilized approximation to a matrix pseudoinverse, commonly used to map Cartesian motions into joint motions near singular configurations. “ denotes the damped pseudoinverse of the Jacobian .”
- Discounted return: The cumulative reward in reinforcement learning after weighting future rewards by a discount factor. “The objective is to maximize the expected discounted return”
- Finite-difference gradient: A derivative estimate formed from function values at nearby points rather than from analytic differentiation. “Increasing the stiffness of these models improves physical fidelity, but makes the dynamics more difficult to resolve accurately with finite simulation timesteps”
- First-order policy optimization: Policy optimization that uses derivatives of the objective with respect to policy parameters. “We integrate BCG into Warp and use it within a SHAC-style actor-critic framework for first-order policy optimization.”
- Gauss–Seidel solver: An iterative numerical method that updates variables sequentially using the latest available values. “contact impulses are computed using a modified Gauss--Seidel solver”
- Gradient propagation: The transmission of derivatives through successive computations, dynamics steps, or layers of a model. “We describe these phases below, followed by their integration into policy learning for motion imitation.”
- Gradient variance: The variability of gradient estimates across samples, environments, or perturbations. “BCG reaches higher final rewards compared to PPO in a few motions as well.”
- Hard-contact model: A contact model that enforces non-penetration and contact constraints directly, typically producing nonsmooth transitions. “Hard-contact models instead express non-penetration and friction through complementarity constraints”
- Interior-point solver: An optimization algorithm that handles inequality constraints by maintaining iterates inside the feasible region. “Dojo uses a Nonlinear Complementarity Problem (NCP) and implicit differentiation of its interior-point solver.”
- Jacobian: A matrix containing the partial derivatives of a vector-valued function with respect to its inputs. “Let and for the transition Jacobians”
- Kinematic chain: A sequence of connected rigid bodies and joints whose configuration determines the position of an end link. “We perturb the contacting kinematic chain in Cartesian space”
- Likelihood-ratio identity: An identity that expresses policy gradients using the derivative of the logarithm of the action probability. “Model-free methods estimate using the likelihood-ratio (score-function) identity”
- Markov decision process (MDP): A sequential decision-making model in which the next state depends probabilistically only on the current state and action. “Policy learning is formalized as a Markov decision process”
- Moreau time-stepping scheme: A nonsmooth numerical integration method for mechanical systems with impacts and contact. “Within Moreau's time-stepping scheme”
- Nonlinear Complementarity Problem (NCP): A problem involving nonlinear functions subject to complementarity conditions, often used to model contact and friction. “Dojo uses a Nonlinear Complementarity Problem (NCP)”
- Principal Component Analysis (PCA): A dimensionality-reduction method that projects data onto directions of greatest variance. “We collect these sensitivities into vectors and use Principal Component Analysis (PCA) to display them in a shared two-dimensional view.”
- Quasi-dynamic: Describing a model that approximates dynamic behavior while simplifying or partially neglecting inertial effects. “enables global planning over quasi-dynamic contact models”
- Quasi-static: Describing a mechanical process assumed to evolve slowly enough that inertial effects are negligible. “These methods operate largely in trajectory optimization or quasi-static manipulation”
- Randomized smoothing: A technique that averages model outputs or derivatives over randomly perturbed inputs to reduce nonsmoothness and variance. “BCG, a contact-local randomized smoothing framework for differentiable policy learning”
- Score-function estimator: A gradient estimator based on the derivative of the log probability of sampled actions. “This estimator differentiates only the policy and treats the dynamics as a black box”
- Sim-to-real transfer: The deployment of a policy trained in simulation onto physical hardware. “We demonstrate zero-shot sim-to-real transfer of dynamic motions learned through differentiable simulation to real humanoid hardware.”
- Sim-to-sim transfer: The evaluation or adaptation of a policy trained in one simulator within another simulator. “We assess sim-to-sim transfer by evaluating the trained policies in MuJoCo.”
- Soft penalty-based contact model: A contact model that represents collisions using compliant forces that increase with penetration. “Soft penalty-based contact models ease differentiation by regularizing contact forces”
- State sensitivity: The derivative of a system state with respect to parameters, inputs, or earlier states. “Let and denote the total sensitivities of the state and action to the policy parameters.”
- Stiff contact: Contact modeled with a large finite stiffness, producing physically realistic interactions but highly sensitive derivatives. “Throughout this paper, stiff contact therefore means a large but finite ”
- Trajectory distribution: The probability distribution over sequences of states and actions induced by a policy. “A stochastic policy with parameters induces a trajectory distribution.”
- Vanishing or exploding gradients: Numerical instability in which propagated derivatives become extremely small or extremely large. “differentiating through long horizons can cause vanishing or exploding gradients”
- Zero-order gradient: A gradient estimate obtained without differentiating through the system dynamics, typically from sampled evaluations. “it is therefore a zero-order gradient that requires no derivatives of the physics”




