---
title: 'MobRT: Digital Twin Framework for Mobile Manipulation'
url: https://www.emergentmind.com/topics/mobrt
type: topic
---

# MobRT: Digital Twin Framework for Mobile Manipulation

Searching arXiv for "MobRT" and closely related papers to ground the article in current literature.
MobRT is a digital twin-based framework for scalable learning in mobile manipulation, introduced to address the difficulty of collecting large-scale, high-quality demonstration data for mobile manipulators that must coordinate base locomotion and arm manipulation in high-dimensional, dynamic, and partially observable environments. It is designed to simulate two primary categories of complex, whole-body tasks—interaction with articulated objects and mobile-base pick-and-place operations—and autonomously generates diverse and realistic demonstrations through the integration of virtual kinematic control and whole-body motion planning [2510.04592].

## 1. Problem setting and scope

The central motivation for MobRT is the disparity between the demands of mobile manipulation and the structure of most existing imitation-learning pipelines. Mobile manipulators require higher-dimensional control than fixed-base tabletop systems, because the policy must control a nonholonomic or car-like mobile base together with a multi-DOF manipulator and gripper. They also require whole-body coordination: for tasks such as opening a door or a dishwasher, the base must move jointly with the end-effector, and separating navigation and manipulation into two planning stages often yields awkward or infeasible motions. These difficulties are compounded by dynamic, cluttered, partially observable environments in which the base changes viewpoint, occlusions vary, and articulated mechanisms have changing kinematic states [2510.04592].

MobRT is positioned against a literature in which most large-scale datasets and benchmarks remain focused on simpler tabletop scenarios. The framework is intended to fill the lack of a digital-twin-based system that can faithfully simulate mobile manipulators, including articulated objects and mobile pick-and-place, and that can automatically generate large-scale, coherent, physically-consistent whole-body trajectories without human teleoperation. Its targeted task families are articulated object interaction, exemplified by opening cabinet drawers and dishwasher doors, and mobile-base pick-and-place, exemplified by picking up a container and placing it on a plate or specified support [2510.04592].

A recurrent misconception in nearby research areas is to treat mobile manipulation as tabletop imitation learning with an additional navigation stage. MobRT is built on the opposite premise: the coupled base-arm configuration space and the need for coordinated motion are primary, not auxiliary, properties of the problem. This design choice shapes both its data-generation stack and its learning benchmark.

## 2. Digital twin construction and task design

MobRT builds a high-fidelity digital twin centered on the ManiSkill3 simulator. The robot mirrors a real mobile manipulator platform described as a Ranger Mini 3-like base plus a Galaxea A1 arm and gripper. The environment comprises indoor scenes with furniture and obstacles; articulated objects from PartNet-Mobility and UniDoorManip; and rigid objects such as containers and plates from RoboTwin-OD together with new AIGC-generated models. Articulated objects retain their joint models, including revolute or prismatic structure, joint limits, and physics, and ManiSkill3 exposes the articulation states directly [2510.04592].

The sensory stack is similarly tied to digital-twin fidelity. MobRT uses multiple RGB-D cameras, specifically two “global” cameras around the base and one wrist camera, mimicking the real setup. Depth rendering uses ray tracing and segmentation to enable realistic visual data and point clouds. Assets are processed with CoACD for convex decomposition, improving collision detection and physically plausible contacts. In the real-world platform described for transfer experiments, the sensor suite is instantiated with three Intel RealSense D435i cameras: two overhead cameras at the back left and right and one wrist-mounted camera [2510.04592].

Task and environment randomization are integral to the framework rather than post hoc augmentation. For each environment reset, the robot base is placed at the world origin, while randomization affects object positions, orientations, and sizes; initial joint configurations of the arm; textures of the ground and selected objects; and lighting conditions. In pick-and-place, objects are scattered on surfaces and targets are placed at random positions. In articulated-object tasks, cabinets and dishwashers appear at different positions and heights, with random initial joint states such as closed or slightly open. This randomization is intended to produce domain diversity and support generalization [2510.04592].

## 3. Automatic demonstration synthesis and whole-body planning

MobRT replaces teleoperation with an automatic demonstration-generation pipeline that maps task specifications into end-effector pose sequences and then into coherent base-plus-arm trajectories. The pipeline begins with digital twin annotation. Each rigid object is annotated with one or more 6-DoF grasp poses, and certain objects such as containers and plates are annotated with a functional axis indicating, for example, “up” for stacking or placing. For mobile pick-and-place, functional-axis alignment defines pre-grasp, grasp, and place keyframes. For articulated objects, MobRT uses Virtual Kinematic Chains (VKC): once the end-effector reaches a handle and the gripper closes, the gripper is treated as rigidly connected to the articulated link, and the articulation model is used to generate an end-effector trajectory from the link trajectory [2510.04592].

For a sampled articulation sequence $\{\theta_t\}_{t=1}^T$, the end-effector trajectory is given by
$$
\mathcal{T}_{\text{eef},t}
=
\mathcal{T}_{\text{link}(\theta_t)}
\,
\mathcal{T}_{\text{link}^{-1}(\theta_{\text{init}})}
\,
\mathcal{T}_{\text{eef}}^{\text{init}},
\quad t=1,\dots,T.
$$
This expresses articulated manipulation as tracking a desired sequence of end-effector poses conditioned on joint evolution rather than as a separate procedural script [2510.04592].

The whole-body planner then solves for a trajectory in the coupled configuration space. With
$$
x(t) = \big[x_{\text{base}}(t),\, q_{\text{arm}}(t)\big], \quad t=1,\dots,T,
$$
MobRT formulates planning as
$$
\begin{aligned}
\min_{x[1:T]} \quad &
\mathcal{C}_{\text{eef}}(x_T, \mathcal{T}_{\text{eef}})
+ \sum_{t=1}^T \Big[\mathcal{C}_{\text{smooth}}(x(t)) + \mathcal{C}_{\text{base}}(x(t)) \Big] \\
\text{s.t.} \quad &
x(1) = x_{\text{init}}, \\
&
x_{\min} \le x(t) \le x_{\max}, \quad \forall t .
\end{aligned}
$$
The end-effector term penalizes pose error, the smoothness term encourages smooth joint and base trajectories, and the base term encodes soft preferences such as approximately fixed chassis orientation, soft collision cost, and minimal unnecessary base motion. After geometric planning, MobRT applies TOPP-RA to compute a time scaling that respects joint velocity and acceleration bounds, yielding dynamically feasible trajectories [2510.04592].

Execution takes place in ManiSkill3 with full physics. Grasping is modeled via rigid connection upon gripper closure; articulated state changes are enforced via the articulation model; and all generated runs are validated. For pick-and-place, success is checked by the distance between the placed object and the target object. For articulated tasks, success is checked by whether the joint angle reaches the target bound. Invalid runs, including motion planning failure, collisions, failed grasp, or insufficient opening, are discarded. The resulting dataset therefore contains only successful and physically valid trajectories [2510.04592].

## 4. Dataset structure, observation-action interface, and policy learning

Each MobRT trajectory records robot state, object state, actions, observations, and task annotations. Robot state includes base pose, arm joint angles, gripper state, and end-effector pose. Object state includes rigid-object poses and articulated joint angles. Observations include RGB images from each camera, depth images fused into point clouds, and segmentation for some procedures such as foreground extraction. Actions consist of the low-level control commands used in planning, serving directly as imitation targets. Task annotations include task identity, object indices, target configuration, and success or failure labels [2510.04592].

The benchmark uses three tasks—Open Cabinet Drawer, Container Place, and Open Dish Washer—with policies trained on **50**, **100**, and **200** demonstration trajectories per task. For sim-to-real, the reported setup uses **300 simulated** demonstrations plus **20 real** demonstrations for Open Cabinet Drawer. The framework is explicitly designed to scale beyond these counts because generation is automated, but the reported experiments focus on performance as a function of demonstration count [2510.04592].

MobRT evaluates four existing visuomotor imitation-learning baselines: ACT, DP, DP3, and iDP3. In parallel, it introduces a multimodal Transformer-based flow-matching policy tailored to mobile manipulators. RGB observations are encoded with a ResNet-18 backbone; depth is converted to point clouds and fused across views; proprioception is encoded separately; and the decoder alternates self-attention over temporal or action tokens with cross-attention to multimodal encoder outputs. Action chunking follows an ALOHA-style formulation, with the policy predicting chunks over a horizon $H$ rather than single controls [2510.04592].

The flow-matching objective is defined over an action chunk $\mathbf{A}$, Gaussian noise $\mathbf{Z} \sim \mathcal{N}(0,I)$, and interpolation time $\tau \in [0,1]$:
$$
\mathbf{X}_\tau = \tau \mathbf{A} + (1-\tau)\mathbf{Z},
$$
with training loss
$$
\mathcal{L}_{\text{FM}}
=
\mathbb{E}_{\tau,\mathbf{Z},\mathbf{A}}
\left[
\left\|
v_\theta(\tau,\mathbf{X}_\tau,o_t) - (\mathbf{A}-\mathbf{Z})
\right\|^2
\right].
$$
Compared to classical diffusion, the framework reports training with 100 steps and inference with around 10 steps, which is presented as more practical for real-time control. Training uses Adam, batch size 64, 100k steps, linear warmup followed by cosine learning-rate decay, and gradient clipping with maximum norm 10 [2510.04592].

## 5. Benchmark results and sim-to-real transfer

A central empirical result is that increasing the number of demonstrations from 50 to 100 to 200 systematically improves success rates across tasks and methods. The paper describes a strong positive correlation between dataset size and task performance, with monotonic improvement and diminishing returns near 200 demonstrations in the reported experiments. This finding is important because it links the digital twin’s automation directly to downstream policy quality rather than treating synthetic trajectory generation as an end in itself [2510.04592].

In direct benchmark comparisons, the proposed MobRT policy outperforms the four baselines. For **Open Dish Washer, 50 demos**, the best baseline is reported at approximately **30%** success, while the proposed RGB-based policy reaches **60%** success. For **Container Place, 100 demos**, baseline DP3 and iDP3 variants are reported at approximately **0–33%** success depending on preprocessing, whereas the proposed point-cloud variant reaches approximately **76.7%** success. The paper also reports that both RGB- and point-cloud-based variants show clear gains over all baselines, with particularly large improvements in low-data regimes [2510.04592].

The benchmark further exposes a representation issue in current 3D diffusion policies. DP3 and iDP3 are described as highly sensitive to background clutter in point clouds: without foreground segmentation they barely improve with more data, while with foreground segmentation, denoted DP3* and iDP3*, success rises sharply. One reported example is **Container Place with 200 demos**, where DP3 increases from **13%** to **73.3%** after foreground segmentation. MobRT attributes its own greater robustness to point-cloud tokenization rather than max-pooling over all points, preserving spatial locality in cluttered scenes [2510.04592].

Real-world transfer is evaluated on the **Open Cabinet Drawer** task using a physical platform with ROS-based control, arm position control, and base velocity control. Two training settings are reported: **20 real demonstrations** only, and **20 real + 300 simulated MobRT demonstrations**. Over **10 real-world rollouts** per policy, iDP3 improves from **10%** success in the real-only setting to **40%** with real-plus-sim data, while the MobRT point-cloud flow-matching policy improves from **20%** to **60%**. The reported failure modes include small steady-state control errors, such as not pulling the drawer far enough, and grasp misalignment under visual conditions that differ substantially from simulation [2510.04592].

## 6. Position in the literature, related usages of the name, and limitations

MobRT is positioned against tabletop-oriented systems such as MimicGen, DexMimicGen, and RoboTwin, which provide sophisticated data generation and high-fidelity digital twins but largely target fixed-base single-arm or dual-arm manipulation. It is also compared with MoMaGen, which is described as the closest framework in spirit for mobile manipulators but one that does not incorporate full whole-body control for complex tasks and remains tightly coupled to replay of human teleoperation trajectories. Other mobile manipulation systems, including Mobile ALOHA, UMI-on-Legs, and TeleMoMa, emphasize teleoperation and data collection, whereas MobRT emphasizes automated simulation-based data generation [2510.04592].

A potential source of ambiguity is the name itself. The explicit titled framework “MobRT” refers to the digital twin-based mobile-manipulation system described above [2510.04592]. In adjacent discussions, however, the label is used more loosely. “Efficient Human-Aware Task Allocation for Multi-Robot Systems in Shared Environments” states that HATA can reasonably be interpreted as a concrete instantiation of a MobRT framework because task-allocation costs and planning are conditioned on human mobility patterns captured in Maps of Dynamics [2508.19731]. “Real-Time Fast Marching Tree for Mobile Robot Motion Planning in Dynamic Environments” similarly describes RT-FMT as essentially a “MobRT” algorithm in the sense of real-time, tree-based motion planning for mobile robots in dynamic environments [2502.09556]. “MRTA-Sim: A Modular Simulator for Multi-Robot Allocation, Planning, and Control in Open-World Environments” presents a simulator explicitly motivated as a testbed for a “MobRT”-type system integrating allocation, planning, and execution [2504.15418]. This suggests an informal broader reading of the term in some contexts, but the named framework on arXiv is the mobile-manipulation system introduced in [2510.04592].

The limitations reported for MobRT are concrete. Task coverage is currently limited to a small set of task types: two articulated categories and one pick-and-place scenario. Physics and perception realism remain incomplete despite depth-noise modeling, domain randomization, and point-cloud preprocessing; the simulator does not capture all real-world effects such as friction variability, actuator backlash, or non-rigid contacts. Policies operate on local observations without explicit high-level planning or language conditioning, and real-world evaluation is concentrated mainly on cabinet drawer opening. The authors identify future directions that include expanding task range and horizon, integrating reinforcement learning to refine skills beyond imitation, improving sim-to-real fidelity, and potentially extending toward multi-robot scenarios or integration with high-level planners and language models [2510.04592].

Source: https://www.emergentmind.com/topics/mobrt