---
title: 'LeHome: Simulator for Deformable Manipulation'
url: https://www.emergentmind.com/topics/lehome
type: topic
---

# LeHome: Simulator for Deformable Manipulation

LeHome is a household robotics simulation environment designed specifically to support manipulation of deformable objects in realistic home settings. Introduced in "LeHome: A Simulation Environment for Deformable Object Manipulation in Household Scenarios" [2604.22363], it addresses a gap in existing household simulators: many prior benchmarks focus on rigid bodies and articulated objects, whereas many real household tasks involve garments, food, fluids, ropes, bags, dust, and similar materials whose shape changes continuously under contact. LeHome combines complete household scenes, deformable-object interactions, multiple robot embodiments, teleoperation support, and a benchmark suite for learning and evaluation, with explicit emphasis on low-cost and open-source robot platforms [2604.22363].

## 1. Problem setting and design goals

The environment is motivated by the observation that household environments are one of the most common, impactful yet challenging application domains for robotics, and that deformable-object manipulation is particularly difficult both in simulation and real-world execution. The paper identifies several reasons: varied categories and shapes, complex dynamics, diverse material properties, multimodal interactions with tools, furniture, and containers, and the lack of reliable deformable-object support in existing simulations [2604.22363].

LeHome frames this difficulty around two underlying bottlenecks. First, collecting large-scale real-world data is expensive and labor-intensive because deformable objects have many possible configurations and home environments are unstructured. Second, accurate simulation is intrinsically hard because one must capture material behavior, nonlinear dynamics, contact, topology changes, and causal action effects. Existing simulators either do not support these objects well or are specialized to a single deformable class, such as only cloth or only soft bodies, and therefore do not provide a broad household benchmark [2604.22363].

The system is organized around three stated goals. The first is broad coverage of deformable household objects rather than specialization to one object family. The second is higher physical realism by selecting simulation methods that match the object type instead of forcing one universal model across all deformables. The third is connection to practical robotic deployment, especially through support for low-cost and open-source robot platforms suitable for household use. The paper repeatedly treats this third point as a central design philosophy rather than a secondary implementation choice [2604.22363].

## 2. System organization and scope

LeHome is structured into three tightly coupled components: **LeHome Assets**, **LeHome Engine**, and **LeHome Benchmark**. LeHome Assets provides the content layer, including household scenes, rigid and articulated assets, deformable objects, and robot embodiments. LeHome Engine is the simulation core, combining different physical simulation methods and interaction mechanisms to model object dynamics and manipulation events. LeHome Benchmark defines representative household tasks and includes teleoperation and data collection tooling. The architecture is presented as a pipeline from realistic scene and object assets, through simulation and interaction logic, into task construction, domain randomization, and demonstration collection [2604.22363].

The environment is intended to cover full household scenes rather than isolated object-only setups. Bedroom, kitchen, living room, and bathroom scenarios are all part of the benchmark design. This full-scene emphasis matters because many deformable household tasks are not only about the object itself but also about interaction with tables, bowls, floors, tools, containers, and articulated furniture. A plausible implication is that LeHome is meant to function as infrastructure for end-to-end embodied evaluation rather than as a narrowly scoped deformable-physics testbed.

In its feature comparison, the paper positions LeHome relative to RoboTwin 2.0, DexGarmentLab, Behavior-1K, Libero, RLBench, and RoboCasa. It is presented as uniquely combining household scenarios, photorealistic rendering, food deformation, flame and particle simulation, fluid simulation, garment manipulation, articulation and rigid manipulation, multi-material manipulation, teleoperation, and support for low-cost robots. The strongest differentiators identified in that comparison are food deformation, flame and particle support, teleoperation, and low-cost robot support [2604.22363].

A common misunderstanding is to reduce LeHome to a garment simulator. The paper explicitly states a broader scope: garments are only one category within a larger household-manipulation agenda that also includes food items, liquids, flame, granules, ropes, posters, bags, and other deformable materials [2604.22363].

## 3. Deformable-object taxonomy and heterogeneous simulation strategy

A core technical idea is the categorization of deformable objects into six classes based on mechanical behavior: **liquid**, **gaseous fluid**, **granular object**, **linear object**, **thin shell**, and **volumetric object**. This taxonomy is operational rather than purely descriptive, because LeHome maps each class to a different physical modeling strategy. The paper explicitly states that it combines PBD, FEM, and Eulerian fluid simulation “to improve physical realism while preserving broad task coverage” [2604.22363].

| Class | Examples | Simulation strategy |
|---|---|---|
| Liquid | water, juice | PBD |
| Gaseous fluid | flame | Omniverse Flow with a sparse voxel grid; Eulerian fluid simulation |
| Granular object | dust, beans, coffee beans | fine granules with PBD; coarse granules approximated as rigid bodies |
| Linear object | cables, ropes | multi-rigid-body chain model or FEM-based deformable model |
| Thin shell | posters, garments, bags | garments with PBD; posters with FEM |
| Volumetric object | patties, sausages, burgers, cutlets | FEM with volumetric discretization |

Liquids are described as materials with no fixed shape but approximately conserved volume, and are simulated with Position-Based Dynamics. The paper argues that PBD is efficient and well suited to frequent contact, which is common in pouring and container interactions. Gaseous fluids are represented mainly by flame; for flame, LeHome uses Omniverse Flow with a sparse voxel grid representation to update key fields, enabling Eulerian fluid simulation with realistic visual appearance [2604.22363].

Granular objects are divided into fine granules and coarse granules. Fine granules such as dust are simulated with PBD to capture dispersed motion, while coarse objects such as coffee beans are approximated as rigid bodies for stable contact interactions. Linear objects such as cables and ropes are modeled either with a multi-rigid-body chain model, which is simpler and more efficient, or with an FEM-based deformable model, which captures more detailed bending and stretching at higher computational cost [2604.22363].

Thin-shell objects include posters, garments, and bags. For highly wrinkling objects like garments, LeHome uses PBD to efficiently simulate stretching and bending. For objects where wrinkling is less important, such as posters, FEM is used for stable elastic response. Volumetric deformable objects such as patties, sausages, burgers, and cutlets are simulated using FEM with volumetric discretization, allowing elastic or elastoplastic stress-strain modeling [2604.22363].

This heterogeneous strategy is one of LeHome’s defining claims. Instead of enforcing a single universal deformable model, the environment chooses a simulator according to the mechanics of the object class and the needs of target household tasks. A common misconception is that “high fidelity” here means a single deeply specified constitutive framework. The paper states the opposite design choice: task-matched heterogeneity. At the same time, it does not provide constitutive law equations, solver settings, timestep information, mesh resolutions, contact model coefficients, or specific material constants, so the fidelity claim is not accompanied by a fully disclosed low-level physics specification [2604.22363].

## 4. Interaction modeling, robots, and data collection pipeline

LeHome’s interaction layer centers on the **Action Graph**, introduced to model cause-effect mechanisms induced by physical interactions, especially those involving morphological splitting or state transitions. The Action Graph uses an event-response structure with three primitives: **attributes**, **nodes**, and **connections**. Attributes carry named data, data types, or metadata; nodes are the computational units and can include trigger nodes, computation nodes, and state update nodes; connections define dataflow and logic dependencies between nodes [2604.22363].

The canonical example is sausage cutting. An On Trigger Node detects collision between knife and sausage and generates a cut-trigger signal; a Computation Node performs mesh segmentation based on a cutting plane; a State Update Node creates new object instances and updates physical properties and textures of the resulting pieces. The intended result is causal consistency: a cut occurs only when triggering conditions are met, after which geometric and state changes follow in an ordered way. The paper presents this as providing modularity, controllability, and extensibility relative to prior approaches [2604.22363].

Robot support spans both mainstream commercial manipulators and low-cost open-source platforms. The introduction mentions UR and Franka as examples of mainstream robots, but the simulator especially emphasizes low-cost platforms from the LeRobot family, including LeRobot, LeKiwi, and XLeRobot. These cover single-arm, bimanual, and mobile-manipulation embodiments. The paper explicitly argues that compact, inexpensive, easy-to-maintain robots are more realistic candidates for widespread home deployment than expensive industrial manipulators, and it frames LeHome as enabling embodied intelligence in “price-sensitive households” [2604.22363].

Teleoperation is another practical component. LeHome supports keyboard and joystick inputs for joint-space commands, a leader-follower system for intuitive joint-synchronized teleoperation, and hybrid control for mobile robots, where the base can be driven by keyboard or joystick while the arms use leader-follower control. The paper emphasizes that the pipeline is compatible with both virtual and real scenarios, allowing simulation datasets to be supplemented with real-world demonstrations under the same workflow [2604.22363].

To narrow the sim-to-real gap, the environment applies domain randomization at the start of each episode over four factors: object initialization positions within feasible workspace bounds, lighting intensity and color temperature, background texture such as tabletop appearance, and visual material properties of objects such as cloth or liquid appearance. The paper states that this visual randomization preserves physical dynamics. After teleoperated data collection, trajectories are replayed under randomized appearance factors to increase visual diversity while keeping object positions fixed so that task geometry and contacts remain valid. Replayed demonstrations are then filtered using task-specific success detectors, such as state validation or geometric constraints, and only successful trajectories are retained [2604.22363].

## 5. Benchmark tasks and policy evaluation

The benchmark consists of six representative household tasks selected to span rooms, interaction types, and manipulation challenges: **Fold Garment** in a bedroom, **Fling Garment** in a bedroom, **Assemble Burger** in a kitchen, **Cut Sausage** in a kitchen, **Pour Coffee** in a living room, and **Wipe Surface** in a bathroom. Collectively, these tasks cover single-arm and bimanual manipulation, tool use, deformable food handling, fluid interaction, rigid object manipulation, rigid-deformable interaction, and deformable-deformable interaction [2604.22363].

The evaluation protocol is explicitly imitation-learning based rather than RL-based. Each policy is trained using **50 teleoperated demonstrations per task**, evaluation is performed over **100 test trials per task**, and the reported metric is **success rate**. The four baselines are **Diffusion Policy (DP)**, an RGB-only imitation policy using action diffusion; **ACT**, a transformer-based imitation policy predicting chunks of actions; **Pi0**, a language-conditioned vision-language-action policy based on flow matching; and **SmolVLA**, a lightweight language-conditioned VLA policy. The paper describes the setup as “unified imitation learning,” using two RGB-only baselines and two language-conditioned VLA baselines to analyze the effect of language conditioning [2604.22363].

The reported success rates are as follows [2604.22363]:

- **Fold Garment**: ACT **45.0%**, DP **30.0%**, SmolVLA **70.0%**, Pi0 **44.0%**.  
- **Assemble Burger**: ACT **78.0%**, DP **35.0%**, SmolVLA **40.0%**, Pi0 **39.0%**.  
- **Fling Garment**: ACT **25.0%**, DP **20.0%**, SmolVLA **36.0%**, Pi0 **14.0%**.  
- **Cut Sausage**: ACT **77.0%**, DP **93.0%**, SmolVLA **90.0%**, Pi0 **75.0%**.  
- **Pour Coffee**: ACT **80.0%**, DP **30.0%**, SmolVLA **40.0%**, Pi0 **40.0%**.  
- **Wipe Surface**: ACT **60.0%**, DP **30.0%**, SmolVLA **60.0%**, Pi0 **60.0%**.  

The authors interpret these results as showing that LeHome can distinguish policy strengths on different household manipulation problems. SmolVLA performs best on garment tasks, Diffusion Policy performs best on **Cut Sausage**, and ACT is competitive on **Assemble Burger** and **Pour Coffee**. The hardest task is clearly **Fling Garment**, where all methods perform poorly. This suggests that large-deformation, long-horizon garment manipulation remains challenging even within the simulator [2604.22363].

## 6. Real-world validation, limitations, and later use

Real-world validation is reported on a limited scale using dual-arm LeRobot hardware. Two training conditions are compared: **Real**, where policies are trained with **10 real-world demonstrations only**, and **Sim+Real Co-Training**, where policies are trained with LeHome simulation demonstrations plus the same **10 real-world demonstrations**. Only ACT and SmolVLA are evaluated, and only on **Fold Garment**, **Assemble Burger**, and **Wipe Surface** [2604.22363].

For ACT, the results are **2/10** to **5/10** on Fold Garment, **2/10** to **4/10** on Assemble Burger, and **1/10** to **7/10** on Wipe Surface when moving from real-only training to co-training. For SmolVLA, the results are **2/10** to **4/10**, **1/10** to **4/10**, and **1/10** to **6/10** on the same tasks. The paper summarizes this as an average improvement from roughly **15%** success to roughly **50%** success, and interprets the result as evidence that LeHome pretraining improves data efficiency and robustness in low-data real-world settings [2604.22363].

At the same time, the paper is explicit or implicit about several limitations. There is no dedicated ablation study on physics model choices, Action Graph design, domain randomization factors, robot embodiment differences, or teleoperation method efficacy. There is no direct quantitative comparison to another simulator on identical tasks, no runtime benchmark, and no explicit measurement of simulator fidelity against ground-truth physical data. Implementation disclosure is partial: the paper names PBD, FEM, Omniverse Flow, Eulerian fluid simulation, sparse voxel grids, mesh segmentation, and volumetric discretization, and it cites NVIDIA Isaac Sim 4.5.0 in the references, but it does not provide solver settings, constitutive parameters, control loop rates, action-vector definitions, or a formal benchmark observation specification [2604.22363].

These omissions matter for interpretation. Another common misunderstanding is to read LeHome’s “high-fidelity” claim as a direct simulator-versus-real calibration result. The paper does not provide direct system identification or simulator-versus-real quantitative error analysis. Its strongest empirical evidence is task outcome performance and limited sim-to-real co-training benefit, not an exhaustive physical validation study [2604.22363].

LeHome also developed into a broader competitive setting. "Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)" [2606.27163] describes the **LeHome Challenge 2026** as an **ICRA 2026 robotics competition** focused on **deformable-object manipulation**, specifically **bimanual garment folding**, evaluated in **Isaac Sim / Isaac Lab**. In that competition, a reinforcement-learning-augmented flow-matching VLA system achieved **1st place out of 62 teams** in the online simulation round with **79.63%** overall success rate and **2nd place** in the real-world final with a score of **865**, indicating that the LeHome ecosystem had already become a platform for method development, competitive benchmarking, and sim-to-real experimentation beyond the original six-task benchmark [2606.27163].

Source: https://www.emergentmind.com/topics/lehome