Brachiation: Locomotion and Robotics
- Brachiation is a dynamic locomotor strategy where animals or robots use arms to suspend, swing, and transfer the body between overhead supports, combining pendulum-like energy exchange and precise grasping and release. Applications include robotic traversal of monkey bars, ladders, and power lines, as well as evolutionary studies of primate locomotion. This strategy is crucial for understanding human evolution and developing advanced robotics, particularly related to swingers and hangers designed for traversing complex and arduous environments.
- Biological brachiation, such as that of gibbons and hypothesized human ancestors, utilises whole-body momentum management and varying arm coordination. Robotic brachiation involves underactuated and hybrid-locomotion dynamics, requiring coordination of nonlinear dynamics, contact transitions, and actuator limits. Successful traversal leverages efficient movement exploitation, energy exchange, and adaptive gripping.
- Model-based and reinforcement learning approaches are used for controlling brachiation in robotic systems. These methods optimize trajectories, handle control authority, and ensure robust and energy-efficient locomotion. They are designed to navigate challenging environments, recover from disturbances, and manage uncertainties and complexities in physical interactions.
Brachiation is a dynamic locomotor strategy in which an animal or robot uses the arms to suspend, swing, and transfer the body between overhead supports such as branches, bars, cables, or handholds. Biological brachiation combines pendulum-like exchange of gravitational potential and kinetic energy, whole-body momentum management, intermittent or continuous contact, and precisely timed grasping and release. In robotics, it is a paradigmatic underactuated and hybrid-locomotion problem: the mechanism generally has fewer actuators than degrees of freedom, while successful traversal depends on coordinating nonlinear dynamics, contact transitions, actuator limits, perception, and recovery from disturbances.
1. Biological forms and evolutionary hypotheses
Gibbons are commonly characterized as specialized one-arm brachiators. In the cross-arm swing described by Fang, Jiang, and Yuan, a gibbon alternates arms, rotates principally about the shoulder, and benefits from a light body, long upper limbs, strong arms, and a large upper-limb/lower-limb length ratio. The proposed human-ancestral alternative, termed “two-arm brachiation,” uses both arms simultaneously to support and propel a heavier body. The animal is hypothesized to hang from a branch with both hands, swing forward, reach or catch another branch with the hands or feet, use the thumbs to assist pushing and gripping, and rotate between vertically extended and forward-inclined postures (Fang et al., 2014).
The authors distinguish this behavior from gibbon cross-arm swinging, chimpanzee-like quadrupedal movement, ordinary climbing, and terrestrial walking or running. In their mechanical interpretation, one-arm brachiation has a shoulder-centered rotational axis, whereas two-arm brachiation rotates more substantially around the lumbar-abdominal region. This difference is used to argue that gibbon-like long arms are not necessarily required for a bilateral mode of suspension.
The proposed sequence includes a bilateral hanging phase, forward body swing, an elevated or “hypsokinetic” posture, reaching or branch contact, thumb-assisted propulsion, and rotation into an “anteverted posture,” described as a forward-bent configuration. These terms are not accompanied by precise joint-angle definitions, motion-capture data, or formal kinematics. The proposal is therefore a qualitative mechanical hypothesis rather than a demonstrated reconstruction of ancestral locomotion.
The evolutionary hypothesis links two-arm brachiation to extension of the knees and hips, lumbar flexibility, longitudinal foot arches, strengthened Achilles tendons, a slimmer body, and a reduced upper-limb/lower-limb ratio. The cited comparative values are approximately $1.4$ for gibbons and $0.85$ for Australopithecus afarensis; modern humans are described as having the smallest ratio among the apes considered. A simplified two-segment model uses
and reports the highest swing frequency near . This result supports only the narrow proposition that a normalized pendulum with similarly sized segments can have favorable timing. It does not establish that Australopithecus used two-arm brachiation or that this behavior selected human limb proportions.
The same evolutionary program has been extended to human shoulder, scapular, thumb, and hunting morphology. Fang and Jiang propose that bilateral suspension may have favored a slim body, more parallel scapulas, mobile shoulders, long and powerful thumbs, and stable grips. These traits were then hypothesized to support throwing, striking, tool control, and hunting after the transition to terrestrial life (Fang et al., 2014). The proposed causal sequence is:
The associated torque argument is expressed as
where is gravitational force, is centrifugal force, and is a moment arm. The equation is dimensionally consistent when the first two quantities are forces and is a distance, but it omits force-angle dependence, muscle counter-torques, scapulothoracic constraints, and tissue adaptation. The inference that repeated loading would drive heritable scapular rotation from oblique to parallel orientation is asserted rather than demonstrated.
The foot-based argument similarly remains speculative. Human-like arches, an aligned hallux, a relatively long foot, and Achilles-tendon anatomy are associated with Australopithecus afarensis, Australopithecus sediba, the A. afarensis fourth metatarsal, and Laetoli footprints. The proposed arboreal pathway involves perpendicular foot orientation during branch jumping, landing, standing, or squatting. Repeated loading in this configuration is hypothesized to have encouraged a longitudinal arch and Achilles-tendon development. The cited fossils and footprints establish early hominin foot features but do not distinguish this pathway from terrestrial bipedalism, climbing, balancing, or mixed locomotor behavior.
2. Mechanics of brachiation
Brachiation is dominated by coupled pendular dynamics. During a supported swing, the body exchanges gravitational potential and kinetic energy while internal joint motion redistributes angular momentum and changes the center of mass. During ricochetal brachiation, one handhold is released, the body follows a ballistic or near-ballistic flight phase, and the next handhold must be grasped within a narrow spatial and temporal tolerance.
A simple radial-loading argument illustrates the mechanical burden of dynamic suspension. For a body of approximately $0.85$0, a swing speed of $0.85$1, and a radius interpreted as $0.85$2, the supporting force can be estimated as
$0.85$3
Using $0.85$4 gives approximately $0.85$5, whereas the cited paper reports approximately $0.85$6. The discrepancy may arise from a different gravitational value, geometric interpretation, or approximation. The total force cannot automatically be assigned to one arm or divided equally between two arms; dynamic asymmetry, joint configuration, branch compliance, and impact forces are not represented.
Several robotic models formalize brachiation with rigid-body manipulator dynamics. A three-link robot can be represented by
$0.85$7
where $0.85$8 is the inertia matrix, $0.85$9 contains Coriolis and centrifugal terms, 0 is gravity, and 1 maps actuator torques into generalized coordinates. The body link changes the center of mass, inertia distribution, and energy exchange during the swing. In a three-link brachiation robot, trajectory optimization through iterative Linear Quadratic Regulator (iLQR) showed the importance of a body link and low-inertia arms for efficient brachiation (Yang et al., 2019).
Brachiation systems are hybrid because continuous swing dynamics are interrupted by discrete release, flight, catch, support-transfer, and coordinate-reset events. A generic transition is
2
where 3 may represent a change in the active support, an impact impulse, or a reset of the support-relative coordinates. In practical systems, release and catch are often handled by passive grippers, heuristics, state machines, or contact constraints rather than a complete compliant contact model.
A four-link model separates upper and lower arms, providing four generalized coordinates and three actuators. Its dynamics are
4
The free gripper is required to reach a target Cartesian position and zero terminal Cartesian velocity:
5
This separates the task-space objective of arriving at a bar from the internal joint-space trajectory used to accomplish it. The four-link structure offers greater obstacle-avoidance freedom than a two-link model, although its supplied validation is simulation-only (Ji et al., 2023).
A minimal single-rod robot instead moves its center of mass along a rigid rod using a crank-slide mechanism. Its generalized coordinates are the rod angle 6 and crank angle 7. The moving mass has radial position
8
with a connecting-rod correction 9. The moving mass modulates both gravitational potential energy and Coriolis-force pumping. In the ideal bang-bang policy, the mass is retracted near stable equilibrium and extended at turning points during swing amplification; during rotation, it is retracted near stable equilibrium and extended near the inverted configuration. The ideal policy is energetically favorable for the model but requires impulsive actuation and instantaneous changes in 0, making it unrealizable.
A continuous input-output-linearizing controller replaces the ideal switching law with critically damped crank motion,
1
where 2. The selected 3 respects an approximate actuator constraint derived from
4
The controller experimentally validated swing amplification and rotation, but not release, aerial transfer, or capture of a subsequent bar (Lieskovský et al., 3 Oct 2025).
3. Robotic morphology and passive support transfer
Robotic brachiators range from single-actuator mechanisms to life-sized humanoids. Their morphologies determine the available momentum-generation mechanisms, the number of controllable coordinates, the complexity of contact transitions, and the tolerance to modeling error.
AcroMonk is a two-arm robot with one central actuator, two articulated arms, and passive unactuated grippers. It has two degrees of freedom while one arm is supported and only one motor torque. Its passive hook gripper has a wide opening, an inclined surface, an off-center connection, and a grooved resting point. The inclined surface is approximately 5, with an overall radius of approximately 6. The gripper slides toward the groove after contact, producing a repeatable support configuration and well-defined start and end poses.
The robot uses an mjbots QDD100 quasi-direct-drive motor with a 7 gear ratio, approximately 8 maximum speed, approximately 9 continuous torque, and approximately 0 peak torque. State feedback combines the motor measurement of relative arm angle with an IMU estimate of the support-arm angle and angular velocity. The deployed controllers run at 1 for PD, 2 for TVLQR, and 3 for RL. AcroMonk demonstrates that passive contact geometry can substitute for active grippers in a strongly underactuated system (Javadi et al., 2023).
RicMonk extends this architecture to a three-link, two-actuator robot weighing approximately 4. Its third link acts as a tail or body-like inertial element. The tail changes the inertia distribution, supports symmetric packaging, reduces wiring and tangling, and enables active redistribution and injection of mechanical energy. RicMonk uses passive conical grooves with larger and smaller radii of approximately 5 and 6, respectively. The conical geometry improves support stability for the heavier robot but increases release torque.
RicMonk supports atomic behaviors designated ZF, ZB, FB, and BF, together with front-release, back-release, front-catch, and back-catch events. A complete cycle includes support, energy-building swing, release, transfer, catch, arm switching, and repetition. The reported release torques are approximately 7 for the support arm and 8 for the swing arm during back release, and approximately 9 and 0 during front release. The robot demonstrated bidirectional brachiation, whereas AcroMonk provided simpler and more robust forward brachiation but could not reliably release in the opposite direction.
For five consecutive forward maneuvers, AcroMonk used 1, had a cost of transport of 2, and required 3. RicMonk used 4, had a cost of transport of 5, and required 6. Cost of transport is defined as
7
Thus, RicMonk consumed more absolute energy but achieved lower mass-normalized transport cost. The comparison illustrates the tradeoff between mechanical simplicity, absolute energy, bidirectionality, and normalized efficiency (Grama et al., 2024).
Passive hooks are also used on a life-sized humanoid for perceptive traversal. The hooks are stainless-steel plates with openings admitting a 8-diameter circle, substantially larger than the 9--0 bar radii used in experiments. Wrist-yaw rotation disengages the hook from the bar without requiring a large arm-lifting motion. The symmetric design supports traversal in either direction and tolerates error in the swing plane (Ongan et al., 30 Aug 2026).
4. Model-based planning and feedback control
Trajectory generation is commonly performed offline because brachiation combines nonlinear dynamics, underactuation, collision constraints, terminal grasp geometry, and free or phase-dependent final times. Direct collocation discretizes the dynamics and jointly optimizes states, torques, duration, terminal conditions, and obstacle avoidance.
For the four-link robot, the reported discrete objective penalizes velocity and torque:
1
The velocity term promotes smooth motion, while the torque term encourages gravitational and passive-dynamic energy exploitation. Collision avoidance represents links as capsules and obstacles as spheres, with minimum squared-distance constraints. The reported simulations used 2 direct-collocation intervals, 3 reference interpolation, less than 4 for direct-collocation optimization, less than 5 for MPC solution, and a 6 control cycle. The energy of actuated joints is measured as
7
Linear MPC tracks both the joint-state trajectory and Cartesian end-effector trajectory. Its objective has the form
8
subject to input bounds. The joint term preserves the planned motion, while the Cartesian term prioritizes accurate arrival at the target bar. In simulation, final end-effector errors were less than 9, optimized-trajectory energy was 0, and MPC-tracking energy was 1. Collision experiments showed that Cartesian feedback could recover the end effector toward its desired path despite substantial joint-space deviations.
Time-varying LQR (TVLQR) stabilizes a nominal trajectory by linearizing the nonlinear dynamics,
2
and solving the finite-horizon Riccati equation
3
The feedback law is
4
TVLQR offers model-dependent feedback and can reduce peak torque, but it is sensitive to mass, inertia, friction, contact, and release-state errors. In RicMonk, TVLQR successfully stabilized ZF and ZB behaviors. BF and FB behaviors required an above-bar approach heuristic to compensate for sim-to-real discrepancies in release-initiated motion.
AcroMonk compares model-free PD, TVLQR, and RL. In isolated BF motion, all three reached 5 success over five trials. Under a cardboard-box disturbance, PD succeeded in 6 trials, TVLQR in 7, and RL in 8. With an additional 9 swing-arm mass, PD compensated reliably, whereas TVLQR and RL failed in 0 tests. Across five consecutive brachiations, PD, TVLQR, and RL used 1, 2, and 3, respectively. RL used the least energy but had a 4 five-swing success rate, compared with 5 for PD and TVLQR. The results show that energy efficiency and robustness to distribution shift are distinct properties.
Robust control for wire-borne brachiators extends the model-based approach to uncertain flexible supports. A two-link robot with one actuator is modeled together with an eight-meter cable. The cable is approximated by three parallel spring-damper elements fitted to the first three vibration harmonics of a higher-fidelity finite-element model. A 6 stiffness uncertainty is included directly in the nonlinear dynamics:
7
The controller uses only the measured angular states, excluding unmeasurable gripper position and velocity. Sum-of-squares (SOS) optimization and semidefinite programming synthesize a polynomial output-feedback controller and a verified inner approximation of a robust backward reachable set. The funnel is
8
The invariance condition requires
9
for all admissible uncertainty values. The formulation explicitly accounts for actuator saturation, unmeasured cable states, parametric stiffness uncertainty, and nonlinear dynamics approximated by polynomials. In simulation and hardware, the SOS controller produced larger verified regions and better off-nominal performance than TVLQR. In a representative uncertain simulation, the SOS controller ended at joint angles 0, with torques between 1 and 2, inside the 3-Nm limit (Farzan et al., 2020).
5. Reinforcement learning and simplified-model imitation
Reinforcement learning treats brachiation as a long-horizon Markov decision process with sparse, timing-sensitive contacts. The primary difficulties are limited control authority, discrete grasp targets, ballistic flight, momentum buildup, and the possibility that a failed grasp is unrecoverable.
A two-stage approach first trains a point-mass model with a virtual extensible arm and then uses its behavior to guide a 14-link articulated model. The simplified model alternates between a spring-damper pendulum hold phase and a ballistic flight phase. Its virtual arm reaches within
4
and its desired arm length is tracked with a maximum force of 5. The simplified policy receives root velocity, swing-phase time, the current handhold, and future handholds, and outputs a target arm-length offset and a binary grab/release flag.
The full model has 14 links and 13 hinge joints, with actuated shoulders, elbows, hips, knees, ankles, and waist, and passive wrists. Its action consists of 13 joint-angle offsets and two grab/release flags. Proximal policy optimization (PPO) trains both simplified and full policies. The simplified model uses approximately 6 parallel processes, while the full model uses approximately 7.
The full-model reward combines task, auxiliary, and style terms:
8
Auxiliary terms track the simplified-model center-of-mass-like trajectory and the next handhold during flight. Style terms penalize excessive torso pitch, arm rotation, knee deviation, and energy use. Directly imposing the simplified model’s release timing degraded performance, whereas providing the reference trajectory as guidance at inference improved it. This suggests that reduced-order timing is useful as a soft constraint but can be harmful when it prevents the full model from adapting to its own morphology and dynamics.
The learned policy exhibits pumping: extra backward and forward swings before release. Pumping emerges to satisfy minimum grasp durations, adjust release angle, and build momentum for large gaps. In a planning mode, the simplified policy evaluates long sequences of candidate handholds, while the full-model value function assesses articulated reachability. A combined reward-plus-value planner passed all three difficult gaps in 9 of $0.85$00 terrains, compared with $0.85$01 for simplified-only planning and $0.85$02 for full-value-only planning (Reda et al., 2022).
A perceptive extension uses Waypoint-Guided Reinforcement Learning (WGRL) on a life-sized dual-arm robot. The method sparsely specifies end-effector waypoints while allowing reinforcement learning to generate whole-body motion. The paper’s available information establishes the use of waypoint guidance, task-success and mechanical-energy rewards, Sim-to-Real training, forward progression, motion stability, geometric variation, hardware evaluation, and failure recovery, but does not provide the robot dimensions, reward coefficients, waypoint representation, policy architecture, or quantitative results (Iwata et al., 18 Aug 2026).
A more complete perceptive system integrates raw lidar, recurrent memory, privileged teachers, and phase-scheduled distillation. An EngineAI PM-01 humanoid uses passive hook end-effectors and a head-mounted RoboSense E1R solid-state lidar. The student consumes a task-specific lidar tensor, proprioception, commands, previous action, and a four-frame history. An attention encoder preserves the two-dimensional scan-grid structure, and a GRU integrates intermittent observations. The student outputs 23 joint-position targets at $0.85$03, while the lidar refreshes at $0.85$04.
Three privileged PPO teachers separately learn jump-up, brachiation, and jump-down. A student is then trained through behavior cloning and DAgger-style data collection, critic warm-up, and PPO refinement with a decaying behavior-cloning anchor. The sim-to-real model includes lidar cone divergence, range noise, edge dropout, sensor mounting perturbations, battery-voltage sag, actuator thermal state, friction, restitution, mass, center-of-mass offsets, PD gains, torque limits, and disturbances.
The complete hardware sequence is
$0.85$05
Across three bar configurations, the system completed $0.85$06 trials, with brachiation speeds up to $0.85$07. The reported success rates were $0.85$08 on Ladder A, $0.85$09 on Ladder B, and $0.85$10 on Ladder C. The only reported failure occurred after successful jump-up when the hook failed to advance to the next bar. The attention-and-memory encoder with an auxiliary centerline loss achieved a behavioral-cloning loss of $0.85$11, compared with $0.85$12 without the auxiliary loss, $0.85$13 for a CNN, $0.85$14 for an MLP, and $0.85$15 for a blind proprioceptive student (Ongan et al., 30 Aug 2026).
6. Estimation, localization, and hybrid contact dynamics
Brachiation is also a demanding setting for state and parameter estimation because support contacts switch, impacts may reset velocities, and payload or support dynamics may be uncertain. A multi-contact estimation framework uses a floating-base humanoid performing monkey-bar maneuvers with an unidentified payload. The estimator jointly infers the state trajectory, initial or arrival state, inertial parameters, and process uncertainties.
For a contact phase $0.85$16, the dynamics are represented as
$0.85$17
with contact transitions represented by reset maps
$0.85$18
The continuous contact dynamics solve accelerations and contact forces simultaneously:
$0.85$19
up to sign convention. During brachiation, the active contact Jacobian changes when a hand is released or a new hand-bar contact is gained.
Each rigid body is represented by a ten-parameter inertial vector containing mass, first mass moment, and rotational inertia. Physical consistency requires nonnegative mass, positive-semidefinite barycentric inertia, and valid principal-moment triangle inequalities. An exponential-eigenvalue parameterization embeds positivity and triangle inequalities directly. A nullspace method handles singularities when principal inertias coincide and their orientation becomes unidentifiable.
The estimator uses multiple shooting and a parametrized Riccati recursion. Unlike ordinary differential dynamic programming, the recursion propagates state curvature, parameter curvature, and state-parameter curvature. Its local policy has the form
$0.85$20
This permits payload parameters to influence future state predictions, contact forces, and measurement residuals. Multiple shooting treats each node state as an independent variable and explicitly recomputes dynamic defects, reducing the sensitivity of long rollouts to erroneous initial states or inertial parameters.
The brachiation example concerns Talos, a humanoid robot performing intricate monkey-bar maneuvers with an unknown payload. The supplied information does not provide Talos’s contact schedule, number of bars, state dimension, inertial ground truth, sampling interval, or quantitative brachiation-specific estimation error. The reported hardware experiment is instead performed on a Go1 quadruped carrying an unknown $0.85$21 payload. That experiment estimated $0.85$22, an absolute error of approximately $0.85$23, or approximately $0.85$24, and therefore validates inertial estimation and localization more generally rather than brachiation hardware performance (Martinez et al., 2024).
7. Applications, limitations, and research directions
Brachiation robots provide platforms for studying underactuation, hybrid contact, passive mechanics, trajectory optimization, robust control, reinforcement learning, perceptive locomotion, and sim-to-real transfer. Applications identified across the research include traversal of monkey bars, ladder-like structures, flexible cables, wires, power lines, branches, and sparse three-dimensional supports. Potential uses include forest exploration, agricultural surveillance, biomimetic locomotion research, and expansion of humanoid traversable workspaces.
Several recurring design principles emerge. A body or tail link can improve energy exchange and momentum generation, while low-inertia arms reduce the effort required for limb repositioning. Passive grippers reduce actuator count and mass and can provide repeatable support poses. Cartesian objectives are valuable because joint-space deviations, particularly in passive coordinates, need not imply failure of end-effector arrival. Robust output-feedback and funnel methods are useful when support states are unmeasurable. Simplified models can guide high-dimensional policies through center-of-mass trajectories, contact timing, and long-horizon handhold planning.
The principal limitations are methodological and physical. Many models are planar, use rigid supports, or omit compliant grasping, frictional contact, branch deformation, impact losses, and hybrid reset dynamics. Direct collocation often treats grasping and releasing at the task or state-machine level rather than through detailed contact mechanics. TVLQR depends on nominal parameters and local linearization. MPC commonly tracks a precomputed collision-free trajectory rather than performing online nonlinear replanning. SOS certificates are formal only relative to reduced polynomial models, finite-degree certificates, bounded uncertainty sets, and time-sampled constraints.
Reinforcement-learning systems face different limitations: prescribed handhold order, structured terrain distributions, partial observability, sensitivity to out-of-distribution states, accumulated release and catch errors, and dependence on reward and termination design. Perceptive systems must cope with thin bars, intermittent lidar returns, sensor noise, actuator thermal limits, battery sag, and sparse geometry. Life-sized systems add substantial whole-body torque, release, and impact demands.
Human-evolutionary interpretations face a separate evidence problem. Comparative anatomy, fossils, footprints, and qualitative mechanics support the presence of relevant anatomical traits and the mechanical plausibility of suspension, but they do not establish that two-arm brachiation caused human bipedalism, scapular orientation, thumb elongation, foot arches, lumbar flexibility, or hunting. The proposed evolutionary pathways remain testable hypotheses requiring comparative primate observations, instrumented brachiation, realistic musculoskeletal and robotic models, fossil-based trait histories, branch-mechanics studies, and quantitative comparisons with climbing, balancing, terrestrial walking, and mixed locomotor repertoires.
Future robotic work includes three-dimensional brachiation, irregular support geometry, online trajectory optimization, automatic release, impact-aware support transfer, short-flight and ricochetal maneuvers, improved passive grippers, moving-horizon estimation, formal safety guarantees, robust perceptive planning, and hardware validation of four-link and single-rod systems. A central technical challenge remains the integration of perception, contact-mode inference, momentum management, actuator constraints, and recovery within a single controller capable of handling both nominal and failed grasp states.
Brachiation is consequently best understood not as one mechanism or one control algorithm, but as a family of dynamic, contact-dependent locomotion problems. Its biological study concerns the relationship between arboreal movement and primate evolution; its robotic study concerns the exploitation and verification of nonlinear underactuated dynamics; and its perceptive-learning study concerns how sparse observations can be converted into reliable whole-body motion across discontinuous supports.