---
title: 'AirExo-3: Robot-Isomorphic Passive Exoskeleton'
url: https://www.emergentmind.com/topics/airexo-3
type: topic
---

# AirExo-3: Robot-Isomorphic Passive Exoskeleton

AirExo-3 is a self-designed robot-isomorphic passive exoskeleton introduced as the hardware front-end of ExoGS, a 4D Real-to-Sim-to-Real framework for scalable manipulation data collection and policy learning. Its purpose is to capture human demonstrations in the real world while preserving robot-consistent kinematics, so that the exoskeleton joint configuration can be interpreted directly as the target robot joint vector without retargeting. In the reported instantiation, AirExo-3 is geometrically matched to a Flexiv Rizon 4s with an Aurora Lite gripper, senses joint motion with encoders only, and is combined with synchronized multi-view RGB-D imaging to support editable 3D Gaussian Splatting reconstruction and geometry-consistent replay of manipulation episodes [2601.18629].

## 1. Functional role and conceptual basis

AirExo-3 was designed to address a specific bottleneck in manipulation learning: obtaining high-quality, robot-consistent manipulation trajectories without using an actual robot during data collection. ExoGS uses it as a robot-free interface through which a human directly performs manipulation with a hand-held gripper while the device records a kinematically consistent trajectory and synchronized RGB observations. The resulting dataset couples action traces with visual observations in a form suitable for subsequent replay, augmentation, and policy learning.

The defining concept is that AirExo-3 is both **robot-isomorphic** and **passive**. Robot-isomorphic means that the exoskeleton’s kinematic chain, including joint types, joint order, link lengths, joint limits, and gripper opening range, is geometrically matched to the target robot’s kinematic model. Passive means that the device has no actuators providing forces to the human; it only senses joint angles with encoders, and all motion is human-driven. This combination is intended to preserve embodiment consistency while keeping the device cheap, light, and easy to maintain.

Because the kinematics are identical, the exoskeleton joint configuration vector is, after calibration, the robot joint vector:
\[
\boldsymbol{q}_t = [q_1,\dots,q_n]^\top,\quad g_t \in [0,1].
\]
A demonstration is stored as
\[
\tau = \{(\boldsymbol{q}_t, g_t)\}_{t=1}^H.
\]
This direct mapping avoids optimization-based retargeting or inverse kinematics during data capture. A common misconception is that AirExo-3 is primarily a teleoperation controller. ExoGS instead positions it as a robot-free acquisition device whose main value lies in collecting physically plausible, robot-valid trajectories before any robot execution occurs.

## 2. Mechanical architecture and kinematic correspondence

The mechanical structure is a serial link-joint chain consisting of seven articulated joints and a parallel gripper. It shares identical kinematic parameters, joint limits, and gripper opening range with the target robot. In the reported system, the target is a Flexiv Rizon 4s with an Aurora Lite gripper. This one-to-one structural correspondence is the basis for the claim that the encoder readings are the robot joint angles after per-joint offset calibration [2601.18629].

All structures are 3D-printed glass-fiber-reinforced nylon, chosen for a lightweight yet stiff and robust construction. Each joint module contains a joint shaft, cross-roller bearings, and a joint body housing a 12-bit miniature rotary encoder. Mechanical limit slots in the shaft and a high-strength pin fixed to the joint body enforce joint limits physically. An optional spring-based suspension at Joint 4 reduces perceived weight, and each joint includes tunable passive damping via a groove and retaining ring for smoother motion and operator comfort.

The kinematic model follows the target robot’s URDF structure. Each exoskeleton joint corresponds to one robot joint:
\[
\text{AirExo joint } i \;\leftrightarrow\; \text{robot joint } i.
\]
Forward kinematics for link \(\ell\) at time \(t\) is written as
\[
\mathbf{T}_{\ell,t} = \mathrm{FK}_\ell(\boldsymbol{q}_t), \quad \ell=1,\dots,L,\;t=1,\dots,H.
\]
The consequence is not merely representational convenience. Mechanical stops mirror robot limits, so the user cannot command poses outside the robot’s admissible joint range, and the reachable workspace of the human hand while using the device matches the robot’s reachable set. ExoGS refers to this property as robot-consistent kinematic constraints.

## 3. Sensing, calibration, and metrological performance

AirExo-3 uses rotary encoders only, with no vision or IMU sensing. There are eight encoders in total: seven for the articulated joints and one for the gripper. The encoders provide 4096 discrete positions per revolution, and the eight devices are connected via a shared bus enabling synchronized joint-state acquisition at up to approximately 300 Hz. The encoder-only design is presented as a deliberate choice: forward-kinematics-based pose recovery is not affected by occlusion, lighting, or camera noise.

Calibration is mechanical and per-joint. Each joint shaft contains a limit slot, and the joint body contains a pin hole. During calibration, an extended pin is inserted so that the joint is mechanically locked at the designed zero angle. The encoder reading \(e_i^{(0)}\) is then recorded for joint \(i\). Subsequent angle computation uses
\[
q_i = \frac{2\pi}{4096} \, \bigl( e_i - e_i^{(0)} \bigr),
\]
with clamping or wrapping to the allowed joint limits defined by the shaft slot [2601.18629].

The reported evaluation adopts an established protocol for exoskeleton-based manipulation systems from AirExo-2 and AirExo, and states that AirExo-3 achieves an average end-effector position error \(< 1\) mm. The error is expressed as
\[
\text{Position error} = \frac{1}{H} \sum_{t=1}^H \left\| \mathbf{p}^{\,\text{exo}}_t - \mathbf{p}^{\,\text{ground truth}}_t \right\|_2.
\]
This millimeter-level accuracy is central to the system’s intended role. It implies that the exoskeleton can provide joint trajectories precise enough for direct use in robot forward kinematics and in subsequent 3D scene reconstruction pipelines, without depending on external marker tracking or motion-capture systems.

## 4. Integration into the ExoGS 4D capture and replay pipeline

AirExo-3 is embedded in a multi-stage pipeline that combines kinematic recording with synchronized multi-view RGB-D imaging. Multiple calibrated RealSense D415 cameras record multi-view images
\[
\mathcal{I} = \{ I_t^{(k)} \}_{t=1,k=1}^{H,K},
\]
with all observations expressed in a common world frame through COLMAP calibration. The exoskeleton provides the joint-angle sequence and gripper opening, and the cameras provide the synchronized visual record of the same interaction [2601.18629].

The scene is reconstructed through a capture-reconstruct-assetize procedure based on editable 3D Gaussian Splatting assets. Camera extrinsics are obtained with COLMAP, and Gaussian parameters are optimized by minimizing a photometric objective of the form
\[
\mathcal{L}_{\text{photo}} = \lambda_1 \,\|\hat{I} - I\|_1 + \lambda_2\, (1-\text{SSIM}(\hat{I}, I)).
\]
The resulting scene is decomposed into independent assets for the robot arm, manipulated objects, and environment/background. This decomposition is essential because replay in simulation is not a monolithic video-based reenactment but a geometry-consistent recomposition of independently transformable assets.

Object pose estimation is performed from the multi-view RGB-D images using FoundationPose. Per-camera object poses \(\mathbf{T}_{o,t}^{(k)} \in SE(3)\) are fused by taking rotation from a primary camera and averaging translations across cameras to obtain \(\{\mathbf{T}_{o,t}\}_{t=1}^H\). Robot link poses are computed directly from the recorded \(\boldsymbol{q}_t\) using the robot URDF:
\[
\mathbf{T}_{\ell,t} = \mathrm{FK}_\ell(\boldsymbol{q}_t), \quad \ell = 1,\dots,L,\; t = 1,\dots,H.
\]

Replay is then constructed by rigidly transforming the robot Gaussians according to the link transforms, transforming object Gaussians by the fused object poses, and keeping environment Gaussians fixed. When an object is grasped, a PoseProcess “fix operation” can rigidly attach the object to the end-effector using a relative transform defined at grasp time. This arrangement means that, because AirExo-3 is robot-isomorphic and calibrated, no extra optimization is needed to align the kinematic model and the rendered 3DGS robot: the kinematic skeleton directly moves the Gaussian assets. A plausible implication is that the hardware’s principal contribution to ExoGS is not only precise motion capture, but the elimination of a substantial alignment problem that would otherwise arise between human demonstration geometry and robot simulation geometry.

## 5. Empirical behavior and comparison with teleoperation

The ExoGS study compares AirExo-3 with a standard robot teleoperation interface using 10 volunteers with approximately 10 minutes of training each. AirExo-3 is reported as consistently faster than teleoperation across three tasks—Pick and Place, Pick Place Close, and Unscrew Bottle Cap—and the inter-user variance is smaller, indicating more consistent performance across users. The qualitative explanation given is that the operator is physically holding the gripper that touches the object, rather than controlling a remote robot, thereby reducing mental and operational burden [2601.18629].

The task-level demonstration success ratios are:

| Task | AirExo-3 | Teleoperation |
|---|---:|---:|
| Pick and place | 100% | 92.3% |
| Pick place close | 100% | 83% |
| Unscrew bottle cap | 87% | 17% |

The difference is especially pronounced for the contact-rich unscrewing task, where teleoperation suffers many failures, collisions, and object damage, while AirExo-3 maintains a substantially higher success rate. This indicates that the device is particularly effective when precise contact sequencing and continuous physical feedback matter more than remote viewpoint control.

The downstream policy results further distinguish raw visual fidelity from trajectory quality. Policies trained directly on teleoperation trajectories may perform better in visually simple domains, but ExoGS policies built from AirExo-3 demonstrations and synthetic augmentation outperform teleoperation-trained policies on the Unscrew Bottle Cap task, achieving \(24\%\) success versus \(8\%\), and on “Pick and place (New Object),” where ExoGS reaches \(76\%\) while teleoperation has \(0\%\). The paper also states that with strong data augmentation and/or Mask Adapter, ExoGS policies eventually outperform teleoperation-trained policies on many generalization scenarios. This suggests that AirExo-3’s main empirical advantage lies not only in faster collection, but in producing cleaner and more transferable contact-rich demonstrations for scalable synthetic-data generation.

## 6. Limitations, design trade-offs, and relation to the broader AirExo line

AirExo-3 inherits several limitations from both its own hardware design and the surrounding ExoGS framework. It is focused on a single serial arm with a simple parallel gripper; multi-arm systems and dexterous hands are not handled. It is passive, so there is no haptic feedback from the robot, and the human supports the weight, albeit partly assisted by the optional spring suspension. Its range of motion is intentionally limited to match the robot, so the operator’s natural reach may be larger than the reachable set available through the device. At the pipeline level, the framework assumes rigid objects, and performance can degrade when segmentation quality drops or when the real environment differs drastically from the simulated rendering conditions [2601.18629].

The design trade-offs are explicit. Passive construction is described as low cost, light, robust, simple, and safe, with a system cost of approximately \(\$400\) including exoskeleton and other hardware, whereas an active exoskeleton could provide force feedback and improved ergonomics at the cost of greater weight, expense, complexity, and safety burden. Encoder-only sensing is robust and accurate, but adding IMUs or external optical trackers could in principle improve some orientation estimates while increasing calibration complexity. Robot-isomorphic constraints eliminate retargeting, but they also force human movement into the robot’s kinematic structure, which may be less natural than freehand manipulation. Mechanical zero-locking yields repeatable calibration, but depends on careful slot-and-pin hardware design.

AirExo-3 is best understood within a broader AirExo lineage. The earlier AirExo system was presented as a low-cost, adaptable, and portable dual-arm exoskeleton for teleoperation and in-the-wild demonstration collection in whole-arm manipulation, with 8 DoF per arm and a two-stage learning framework that combined a small amount of teleoperated robot data with extensive in-the-wild exoskeleton demonstrations [2309.14975]. AirExo-3 differs in role and emphasis: it is a self-designed robot-isomorphic passive exoskeleton used as the capture interface inside ExoGS, tightly coupled to 3D Gaussian Splatting reconstruction, multi-view RGB-D observation, and Real-to-Sim-to-Real policy learning. This suggests a specialization from broad whole-arm imitation toward high-fidelity, robot-consistent data acquisition for geometry-consistent replay and scalable synthetic manipulation data generation.

The ExoGS authors state that code and hardware files have been released at `https://github.com/zaixiabalala/ExoGS`. In that sense, AirExo-3 is not only a device but a reproducible hardware-software component in a larger experimental stack whose central claim is that high-quality robot-consistent motion capture can be achieved without collecting demonstrations on the robot itself.

Source: https://www.emergentmind.com/topics/airexo-3