---
title: ICRA 2024 Cloth Competition
url: https://www.emergentmind.com/topics/icra-2024-cloth-competition
type: topic
---

# ICRA 2024 Cloth Competition

The ICRA 2024 Cloth Competition was a head-to-head benchmark for **grasp pose selection for in-air robotic cloth unfolding**, created to address the lack of standardized benchmarks and shared datasets in robotic cloth manipulation. It combined a fixed dual-arm hardware platform, a single-attempt unfolding protocol, a public dataset of real-world robotic trials, and a live on-site evaluation at ICRA. In its published form, the benchmark comprises **679 unfolding demonstrations across 34 garments**, including **176 competition evaluation trials** recorded during the event, and it emphasizes independent, out-of-the-lab evaluation of cloth grasp-selection methods [2508.16749].

## 1. Origins and benchmark rationale

Robotic cloth manipulation has been difficult to benchmark because cloth is highly deformable, self-occluding, and sensitive to small variations in initial configuration, material, and gravity-driven dynamics. The competition was organized to address three linked problems: reproducibility is difficult because the same garment rarely appears in exactly the same configuration; comparison across methods is hard because prior papers often use different setups, garments, and evaluation metrics; and data-driven methods remain data-starved because collecting labeled real-world interaction data is expensive and non-trivial [2508.16749].

The benchmark therefore defines a deliberately narrow scope: **single in-air regrasp + stretch**. This isolates one key problem in cloth manipulation, namely **grasp pose selection on hanging cloth for unfolding**, while standardizing hardware, camera viewpoint, and evaluation metrics. A plausible implication is that the organizers prioritized methodological comparability over full-system autonomy. That design choice distinguishes the competition from broader cloth-manipulation pipelines that mix initialization, perception, motion planning, and recovery behaviors in different proportions [2508.16749].

The competition also sits within a longer benchmarking trajectory in deformable-object manipulation. The "Household Cloth Object Set" proposed a set of household cloth objects and related tasks as a roadmap toward common benchmarks in cloth manipulation [2111.01527]. In the perception subdomain, CeDiRNet-3DoF later reported first place in the perception task of ICRA 2023's Cloth Manipulation Challenge and emphasized the lack of standardized benchmarks for cloth grasping [2408.14456]. A later real-world evaluation framework, DRAPER, likewise focused on fair comparison under shared fabrics, protocols, and metrics for flattening and folding [2409.15159]. This suggests that the ICRA 2024 competition should be understood not as an isolated event but as part of a broader effort to make cloth-manipulation results comparable across laboratories and methods.

## 2. Task formulation and standardized physical setup

The competition benchmark models **in-air unfolding via grasp pose selection**. A garment lies crumpled on a table; the robot first picks it up and makes it hang from two points; an algorithm then chooses a new grasp pose on the hanging cloth; and the robot executes a final grasp-and-stretch in the air to unfold it. In this benchmark, grasp pose selection means selecting one or more grasp poses for the **right arm**, with each pose represented as a **6-DoF frame in robot/world coordinates**: position $\mathbf{p} \in \mathbb{R}^3$ and orientation $R \in SO(3)$, stored as a $4 \times 4$ homogeneous transform in JSON [2508.16749].

The full unfolding procedure is fixed. The **right UR5e** autonomously grasps the highest visible point on the crumpled garment on the table and lifts it. The **left UR5e** then autonomously grasps the lowest point of the hanging cloth. After this initialization, the system captures a start observation of the hanging garment, the participant’s algorithm receives that observation and returns one or more candidate grasp poses for the right gripper, the first reachable pose is executed through a pregrasp and linear approach, and both arms move to stretch the cloth until tension reaches **2 N** via wrist force-torque sensors, or until a grasp failure is detected [2508.16749].

The hardware is standardized across dataset collection and on-site evaluation. The setup uses **two UR5e robot arms, 90 cm apart**, each equipped with a **Robotiq 2F-85 parallel gripper** with **85 mm stroke** and **rubber-coated fingertips**. Sensing is provided by a **ZED 2i stereo RGB-D camera**, mounted **80 cm above and 140 cm behind the robots** at an approximately **20° downward angle**, with **neural depth mode via ZED SDK 4.0**. The world frame is centered between the robot bases, and robot base poses, tool poses, and camera intrinsics/extrinsics are recorded for each observation [2508.16749].

An important feature of the interface is that participants do **not** handle motion planning or collision checking. The benchmark infrastructure computes a pregrasp pose by translating a few centimeters along the negative gripper $z$-axis, moves the arm collision-free to that pregrasp pose, then moves linearly to the final grasp pose and closes the gripper. The algorithmic responsibility is restricted to outputting grasp poses within the allotted time budget [2508.16749].

## 3. Dataset composition and data representation

The full dataset associated with the competition contains **679 unfolding demonstrations** across **34 garments**. It is composed of **503 training demonstrations** collected in the AIRO lab and **176 competition evaluation trials** recorded live at ICRA 2024. The training set consists of human-annotated grasp poses intended for unfolding, whereas the competition subset records the actual performance of the participating teams on the shared physical platform [2508.16749].

Garment coverage is structured but diverse. The training dataset uses **30 cloth items**: **15 T-shirts** and **15 towels**, with the towels including **10 items from the Household Cloth Object Set** [2508.16749]. The competition itself uses **8 garments**: **4 T-shirts** and **4 towels**. Of these, **4 were present in the training dataset** and **4 were new**, enabling analysis of generalization to unseen garments under live evaluation conditions [2508.16749].

Each episode contains a **start observation folder**, a **grasp folder**, a **result observation folder**, and an **interaction video**. The start and result observation folders include `image_left.png` and `image_right.png` stereo RGB images at **2208×1242×3** `uint8`, `depth_map.tiff` aligned with the left view at **2208×1242** `float32`, `confidence_map.tiff`, `depth_image.jpg`, and a colored `point_cloud.ply` in world coordinates with approximately **2.7M points**. Robot states are stored in JSON files for left and right joints, base poses, and TCP poses, and camera parameters include `camera_intrinsics.json`, `camera_pose_in_world.json`, and `right_camera_pose_in_left_camera.json` [2508.16749].

Pose representations follow standard robotics notation:
$$
T \in \mathbb{R}^{4 \times 4}, \qquad
T =
\begin{bmatrix}
R & \mathbf{p} \\
0 & 1
\end{bmatrix},
$$
with $R \in SO(3)$ and $\mathbf{p} \in \mathbb{R}^3$ [2508.16749]. The grasp folder stores the selected grasp pose, typically as such a transform in world coordinates. In the human-annotated training subset, annotators choose a point on the point cloud and set orientation tangent to the cloth surface [2508.16749].

A **reference observation** is provided per garment, showing the cloth fully unfolded and stretched under ideal conditions. This reference is used to compute the maximum projected area for coverage evaluation. All data are stored in standard formats—PNG, TIFF, PLY, JSON, and MP4—and Python loading utilities and notebooks are provided for visualization, segmentation, and conversion workflows [2508.16749].

## 4. Evaluation protocol and scoring

For each trial, the algorithm receives the full start observation: stereo RGB images, depth map, colored point cloud, robot poses and joint states, and camera calibration. It must return one or more grasp poses for the right gripper in the world frame, represented as a list $\{T_k\}_{k=1}^K$ of candidate $4 \times 4$ transforms. The benchmark system tries these sequentially until it finds a reachable pose. A valid submission is a program, typically in Python, that accepts the observation via the competition server and returns grasp poses within **30 s per trial** [2508.16749].

Three quantitative descriptors structure evaluation. Let $A_{\text{curr}}$ be the projected area of the garment in the final stretched observation and $A_{\text{max}}$ the maximum projected area from the garment-specific reference observation. Then **coverage** is
$$
\text{coverage} = \frac{A_{\text{curr}}}{A_{\text{max}}} \in [0,1].
$$
Grasp success is binary:
$$
\text{success}(i) =
\begin{cases}
1, & \text{if cloth is present between fingertips after grasp} \\
0, & \text{otherwise}.
\end{cases}
$$
Over $N$ trials, **grasp success rate** is
$$
\text{SuccessRate} = \frac{1}{N} \sum_{i=1}^{N} \text{success}(i),
$$
and **coverage for successful grasps** is
$$
\text{Coverage}_{\text{success}} =
\frac{1}{\sum_{i=1}^{N} \text{success}(i)}
\sum_{i=1}^{N} \text{success}(i)\,\text{coverage}(i).
$$
The competition ranking itself uses **mean coverage over all 16 trials**, including failures [2508.16749].

The on-site evaluation protocol is rigid. **Eleven teams** participated. Each team performed **2 trials per garment** on **8 garments**, for **16 trials total**, within a **50-minute timeslot**. The organizers’ server sent the start observation over a local network, the team’s code ran on its own laptop or workstation, and the dual-arm UR5e system executed the submitted grasp and stretch autonomously. This design removed hardware variation from the scoring loop and made the evaluation genuinely shared rather than laboratory-specific [2508.16749].

A common misconception is that cloth-unfolding benchmarks primarily test motion planning or force control. Here the benchmark explicitly tests **grasp pose selection** under a fixed downstream execution pipeline. Because participants do not control the initialization procedure, the stretch trajectory, or the collision checker, performance differences are attributable mainly to the quality of the predicted grasp pose distribution rather than to end-to-end system engineering [2508.16749].

## 5. Participating methods and empirical outcomes

The **11 participating methods** split into **2 traditional geometric methods** and **9 learning-based methods**. The traditional entries were **Intuitive Grasping Determination (IGD, AIR-jnu)** and **Sharp Edge Detection (Ewha Glab)**. The learning-based methods were grouped by supervision type: **graspability estimation**, **semantic keypoint detection**, **grasp imitation**, and **coverage prediction** [2508.16749].

IGD, the winning method, segments cloth in RGB, fits a low-resolution polygon to its boundary, and selects the **second furthest vertex** from the TCP of the left holding gripper as the grasp point, using a fixed grasp orientation. Sharp Edge Detection applies sharp feature detection on the point cloud, identifies prominent sharp points as candidates, selects the candidate closest to the camera, and sets grasp orientation using local surface normals. These two methods are explicitly **class-agnostic**, do not use cloth category or semantic keypoints, and rely solely on geometry [2508.16749].

Learning-based submissions were more heterogeneous. **KeypointDetr** used an end-to-end transformer-like detector to classify garment type and detect semantic keypoints; **CeDiRNet-6DoF** extended center-direction regression to predict 6-DoF grasp poses and combined pre-training on an additional towel dataset with fine-tuning on the competition data; **CopGNN** predicted coverage from a graph built on the point cloud; **Depth2Grasp-CNN**, **CFAN**, and **PointNet-VAE** imitated human grasp selection from depth or point-cloud inputs; and **AED** and **Grasp-Cloth** emphasized graspability rather than direct unfolding outcome prediction [2508.16749].

The top three teams illustrate the central trade-off between grasp success and coverage.

| Team / method | SuccessRate | Overall coverage |
|---|---:|---:|
| AIR-jnu / IGD | $\approx 0.69$ | $\approx 0.60$ |
| Team Ljubljana / CeDiRNet-6DoF | $\approx 0.63$ | $\approx 0.57$ |
| Ewha Glab / Sharp Edge Detection | $\approx 0.94$ | $\approx 0.55$ |

The accompanying **Coverage\_success** values clarify the ranking logic: IGD achieved **$\approx 0.73$**, CeDiRNet-6DoF **$\approx 0.69$**, and Sharp Edge Detection **$\approx 0.57$** [2508.16749]. Methods such as **AED** and **PointNet-VAE** were high-risk/high-reward: they had very low success rates, **0.19** and **0.06**, but high coverage when they did succeed, **0.75** and **0.91**, respectively [2508.16749].

Across all **176 trials**, shirts were easier to grasp than towels: **grasp success rate** was approximately **0.70** for shirts and **0.53** for towels, while **Coverage\_success** was approximately **0.60** for shirts and **0.57** for towels. Seen and unseen garments showed little separation in aggregate performance: average coverage was approximately **0.49** on seen garments and **0.47** on unseen garments, with no significant degradation reported for unseen items [2508.16749].

One of the most discussed outcomes was that **traditional non-learning methods achieved first and third place**, while a learning-based method took second. This does not imply that learned methods were ineffective; rather, it showed that simple geometric reasoning with well-chosen heuristics remained highly competitive under the competition’s single-attempt, failure-inclusive scoring rule. A plausible implication is that the evaluation favored methods that balanced ambitious unfolding with conservative grasp reliability, rather than maximizing only best-case coverage [2508.16749].

## 6. Interpretation, limitations, and legacy

The published analysis emphasized a substantial discrepancy between competition performance and much of the prior literature. Previous works such as **FlingBot** and **UnFoldIR** reported coverage around **0.8** and **0.85**, but they often used multiple actions, filtered out failed grasps, reset between attempts, or evaluated in more controlled, lab-specific setups. By contrast, the ICRA 2024 competition enforced a **single grasp + stretch** attempt, included grasp failures in the score, and evaluated all teams on the same physical platform. The resulting lower top coverage, approximately **0.60**, was അവതരിപ്പated not as regression but as a more realistic and stringent measure of performance [2508.16749].

Several limitations were acknowledged. The dataset was collected on a single **dual-UR5e + ZED2i** platform, so viewpoint and kinematic biases remain. The **single-attempt constraint** does not capture the full potential of iterative unfolding strategies. The dataset contains no explicit cloth-category or semantic-keypoint annotations, although teams can add their own. Some training videos are incomplete because of recording issues, though the **176 competition videos are complete**, and the core observation data remain intact [2508.16749].

The benchmark nevertheless provides a concrete foundation for subsequent work. Later real-world comparison frameworks such as DRAPER defined standardized metrics such as **Normalised Coverage**, **Normalised Improvement**, and **Intersection-over-Union**, together with shared robot and fabric protocols for flattening and folding [2409.15159]. This suggests a convergence toward common cloth-evaluation principles: fixed task protocols, standardized sensing and gripper assumptions, and explicit accounting for failures. In parallel, perception-focused work such as CeDiRNet-3DoF and the ViCoS Towel Dataset showed that competition-aligned public datasets can substantially improve comparability in grasp-point localization [2408.14456].

Within the history of cloth benchmarking, the ICRA 2024 Cloth Competition therefore occupies a specific position. It is not a general benchmark for all deformable-manipulation skills, nor a complete folding or dressing suite. Rather, it is a narrowly scoped, physically grounded benchmark for **grasp pose selection on hanging cloth for unfolding**, backed by a public interaction dataset and live independent evaluation. Its main contribution is methodological: it converts a historically lab-specific subproblem into a reproducible benchmark with shared data, shared hardware, and directly comparable results, and it does so in a form that the published work explicitly proposes as a foundation for future benchmarks and further progress in data-driven robotic cloth manipulation [2508.16749].

Source: https://www.emergentmind.com/topics/icra-2024-cloth-competition