Papers
Topics
Authors
Recent
Search
2000 character limit reached

ICRA 2024 Cloth Competition

Updated 9 July 2026
  • The ICRA 2024 Cloth Competition is a benchmark that standardizes grasp pose selection for in-air unfolding using a dual-arm UR5e platform.
  • It features a public dataset of 679 demonstrations across 34 garments, enforcing a single-attempt protocol to ensure reproducibility.
  • The competition compares traditional geometric and learning-based methods, highlighting trade-offs in grasp success rate and coverage.

The ICRA 2024 Cloth Competition was a head-to-head benchmark for grasp pose selection for in-air robotic cloth unfolding, created to address the lack of standardized benchmarks and shared datasets in robotic cloth manipulation. It combined a fixed dual-arm hardware platform, a single-attempt unfolding protocol, a public dataset of real-world robotic trials, and a live on-site evaluation at ICRA. In its published form, the benchmark comprises 679 unfolding demonstrations across 34 garments, including 176 competition evaluation trials recorded during the event, and it emphasizes independent, out-of-the-lab evaluation of cloth grasp-selection methods (Gusseme et al., 22 Aug 2025).

1. Origins and benchmark rationale

Robotic cloth manipulation has been difficult to benchmark because cloth is highly deformable, self-occluding, and sensitive to small variations in initial configuration, material, and gravity-driven dynamics. The competition was organized to address three linked problems: reproducibility is difficult because the same garment rarely appears in exactly the same configuration; comparison across methods is hard because prior papers often use different setups, garments, and evaluation metrics; and data-driven methods remain data-starved because collecting labeled real-world interaction data is expensive and non-trivial (Gusseme et al., 22 Aug 2025).

The benchmark therefore defines a deliberately narrow scope: single in-air regrasp + stretch. This isolates one key problem in cloth manipulation, namely grasp pose selection on hanging cloth for unfolding, while standardizing hardware, camera viewpoint, and evaluation metrics. A plausible implication is that the organizers prioritized methodological comparability over full-system autonomy. That design choice distinguishes the competition from broader cloth-manipulation pipelines that mix initialization, perception, motion planning, and recovery behaviors in different proportions (Gusseme et al., 22 Aug 2025).

The competition also sits within a longer benchmarking trajectory in deformable-object manipulation. The "Household Cloth Object Set" proposed a set of household cloth objects and related tasks as a roadmap toward common benchmarks in cloth manipulation (Garcia-Camacho et al., 2021). In the perception subdomain, CeDiRNet-3DoF later reported first place in the perception task of ICRA 2023's Cloth Manipulation Challenge and emphasized the lack of standardized benchmarks for cloth grasping (Tabernik et al., 2024). A later real-world evaluation framework, DRAPER, likewise focused on fair comparison under shared fabrics, protocols, and metrics for flattening and folding (Kadi et al., 2024). This suggests that the ICRA 2024 competition should be understood not as an isolated event but as part of a broader effort to make cloth-manipulation results comparable across laboratories and methods.

2. Task formulation and standardized physical setup

The competition benchmark models in-air unfolding via grasp pose selection. A garment lies crumpled on a table; the robot first picks it up and makes it hang from two points; an algorithm then chooses a new grasp pose on the hanging cloth; and the robot executes a final grasp-and-stretch in the air to unfold it. In this benchmark, grasp pose selection means selecting one or more grasp poses for the right arm, with each pose represented as a 6-DoF frame in robot/world coordinates: position p∈R3\mathbf{p} \in \mathbb{R}^3 and orientation R∈SO(3)R \in SO(3), stored as a 4×44 \times 4 homogeneous transform in JSON (Gusseme et al., 22 Aug 2025).

The full unfolding procedure is fixed. The right UR5e autonomously grasps the highest visible point on the crumpled garment on the table and lifts it. The left UR5e then autonomously grasps the lowest point of the hanging cloth. After this initialization, the system captures a start observation of the hanging garment, the participant’s algorithm receives that observation and returns one or more candidate grasp poses for the right gripper, the first reachable pose is executed through a pregrasp and linear approach, and both arms move to stretch the cloth until tension reaches 2 N via wrist force-torque sensors, or until a grasp failure is detected (Gusseme et al., 22 Aug 2025).

The hardware is standardized across dataset collection and on-site evaluation. The setup uses two UR5e robot arms, 90 cm apart, each equipped with a Robotiq 2F-85 parallel gripper with 85 mm stroke and rubber-coated fingertips. Sensing is provided by a ZED 2i stereo RGB-D camera, mounted 80 cm above and 140 cm behind the robots at an approximately 20° downward angle, with neural depth mode via ZED SDK 4.0. The world frame is centered between the robot bases, and robot base poses, tool poses, and camera intrinsics/extrinsics are recorded for each observation (Gusseme et al., 22 Aug 2025).

An important feature of the interface is that participants do not handle motion planning or collision checking. The benchmark infrastructure computes a pregrasp pose by translating a few centimeters along the negative gripper zz-axis, moves the arm collision-free to that pregrasp pose, then moves linearly to the final grasp pose and closes the gripper. The algorithmic responsibility is restricted to outputting grasp poses within the allotted time budget (Gusseme et al., 22 Aug 2025).

3. Dataset composition and data representation

The full dataset associated with the competition contains 679 unfolding demonstrations across 34 garments. It is composed of 503 training demonstrations collected in the AIRO lab and 176 competition evaluation trials recorded live at ICRA 2024. The training set consists of human-annotated grasp poses intended for unfolding, whereas the competition subset records the actual performance of the participating teams on the shared physical platform (Gusseme et al., 22 Aug 2025).

Garment coverage is structured but diverse. The training dataset uses 30 cloth items: 15 T-shirts and 15 towels, with the towels including 10 items from the Household Cloth Object Set (Gusseme et al., 22 Aug 2025). The competition itself uses 8 garments: 4 T-shirts and 4 towels. Of these, 4 were present in the training dataset and 4 were new, enabling analysis of generalization to unseen garments under live evaluation conditions (Gusseme et al., 22 Aug 2025).

Each episode contains a start observation folder, a grasp folder, a result observation folder, and an interaction video. The start and result observation folders include image_left.png and image_right.png stereo RGB images at 2208×1242×3 uint8, depth_map.tiff aligned with the left view at 2208×1242 float32, confidence_map.tiff, depth_image.jpg, and a colored point_cloud.ply in world coordinates with approximately 2.7M points. Robot states are stored in JSON files for left and right joints, base poses, and TCP poses, and camera parameters include camera_intrinsics.json, camera_pose_in_world.json, and right_camera_pose_in_left_camera.json (Gusseme et al., 22 Aug 2025).

Pose representations follow standard robotics notation:

T∈R4×4,T=[Rp 01],T \in \mathbb{R}^{4 \times 4}, \qquad T = \begin{bmatrix} R & \mathbf{p} \ 0 & 1 \end{bmatrix},

with R∈SO(3)R \in SO(3) and p∈R3\mathbf{p} \in \mathbb{R}^3 (Gusseme et al., 22 Aug 2025). The grasp folder stores the selected grasp pose, typically as such a transform in world coordinates. In the human-annotated training subset, annotators choose a point on the point cloud and set orientation tangent to the cloth surface (Gusseme et al., 22 Aug 2025).

A reference observation is provided per garment, showing the cloth fully unfolded and stretched under ideal conditions. This reference is used to compute the maximum projected area for coverage evaluation. All data are stored in standard formats—PNG, TIFF, PLY, JSON, and MP4—and Python loading utilities and notebooks are provided for visualization, segmentation, and conversion workflows (Gusseme et al., 22 Aug 2025).

4. Evaluation protocol and scoring

For each trial, the algorithm receives the full start observation: stereo RGB images, depth map, colored point cloud, robot poses and joint states, and camera calibration. It must return one or more grasp poses for the right gripper in the world frame, represented as a list {Tk}k=1K\{T_k\}_{k=1}^K of candidate 4×44 \times 4 transforms. The benchmark system tries these sequentially until it finds a reachable pose. A valid submission is a program, typically in Python, that accepts the observation via the competition server and returns grasp poses within 30 s per trial (Gusseme et al., 22 Aug 2025).

Three quantitative descriptors structure evaluation. Let AcurrA_{\text{curr}} be the projected area of the garment in the final stretched observation and R∈SO(3)R \in SO(3)0 the maximum projected area from the garment-specific reference observation. Then coverage is

R∈SO(3)R \in SO(3)1

Grasp success is binary:

R∈SO(3)R \in SO(3)2

Over R∈SO(3)R \in SO(3)3 trials, grasp success rate is

R∈SO(3)R \in SO(3)4

and coverage for successful grasps is

R∈SO(3)R \in SO(3)5

The competition ranking itself uses mean coverage over all 16 trials, including failures (Gusseme et al., 22 Aug 2025).

The on-site evaluation protocol is rigid. Eleven teams participated. Each team performed 2 trials per garment on 8 garments, for 16 trials total, within a 50-minute timeslot. The organizers’ server sent the start observation over a local network, the team’s code ran on its own laptop or workstation, and the dual-arm UR5e system executed the submitted grasp and stretch autonomously. This design removed hardware variation from the scoring loop and made the evaluation genuinely shared rather than laboratory-specific (Gusseme et al., 22 Aug 2025).

A common misconception is that cloth-unfolding benchmarks primarily test motion planning or force control. Here the benchmark explicitly tests grasp pose selection under a fixed downstream execution pipeline. Because participants do not control the initialization procedure, the stretch trajectory, or the collision checker, performance differences are attributable mainly to the quality of the predicted grasp pose distribution rather than to end-to-end system engineering (Gusseme et al., 22 Aug 2025).

5. Participating methods and empirical outcomes

The 11 participating methods split into 2 traditional geometric methods and 9 learning-based methods. The traditional entries were Intuitive Grasping Determination (IGD, AIR-jnu) and Sharp Edge Detection (Ewha Glab). The learning-based methods were grouped by supervision type: graspability estimation, semantic keypoint detection, grasp imitation, and coverage prediction (Gusseme et al., 22 Aug 2025).

IGD, the winning method, segments cloth in RGB, fits a low-resolution polygon to its boundary, and selects the second furthest vertex from the TCP of the left holding gripper as the grasp point, using a fixed grasp orientation. Sharp Edge Detection applies sharp feature detection on the point cloud, identifies prominent sharp points as candidates, selects the candidate closest to the camera, and sets grasp orientation using local surface normals. These two methods are explicitly class-agnostic, do not use cloth category or semantic keypoints, and rely solely on geometry (Gusseme et al., 22 Aug 2025).

Learning-based submissions were more heterogeneous. KeypointDetr used an end-to-end transformer-like detector to classify garment type and detect semantic keypoints; CeDiRNet-6DoF extended center-direction regression to predict 6-DoF grasp poses and combined pre-training on an additional towel dataset with fine-tuning on the competition data; CopGNN predicted coverage from a graph built on the point cloud; Depth2Grasp-CNN, CFAN, and PointNet-VAE imitated human grasp selection from depth or point-cloud inputs; and AED and Grasp-Cloth emphasized graspability rather than direct unfolding outcome prediction (Gusseme et al., 22 Aug 2025).

The top three teams illustrate the central trade-off between grasp success and coverage.

Team / method SuccessRate Overall coverage
AIR-jnu / IGD R∈SO(3)R \in SO(3)6 R∈SO(3)R \in SO(3)7
Team Ljubljana / CeDiRNet-6DoF R∈SO(3)R \in SO(3)8 R∈SO(3)R \in SO(3)9
Ewha Glab / Sharp Edge Detection 4×44 \times 40 4×44 \times 41

The accompanying Coverage_success values clarify the ranking logic: IGD achieved 4×44 \times 42, CeDiRNet-6DoF 4×44 \times 43, and Sharp Edge Detection 4×44 \times 44 (Gusseme et al., 22 Aug 2025). Methods such as AED and PointNet-VAE were high-risk/high-reward: they had very low success rates, 0.19 and 0.06, but high coverage when they did succeed, 0.75 and 0.91, respectively (Gusseme et al., 22 Aug 2025).

Across all 176 trials, shirts were easier to grasp than towels: grasp success rate was approximately 0.70 for shirts and 0.53 for towels, while Coverage_success was approximately 0.60 for shirts and 0.57 for towels. Seen and unseen garments showed little separation in aggregate performance: average coverage was approximately 0.49 on seen garments and 0.47 on unseen garments, with no significant degradation reported for unseen items (Gusseme et al., 22 Aug 2025).

One of the most discussed outcomes was that traditional non-learning methods achieved first and third place, while a learning-based method took second. This does not imply that learned methods were ineffective; rather, it showed that simple geometric reasoning with well-chosen heuristics remained highly competitive under the competition’s single-attempt, failure-inclusive scoring rule. A plausible implication is that the evaluation favored methods that balanced ambitious unfolding with conservative grasp reliability, rather than maximizing only best-case coverage (Gusseme et al., 22 Aug 2025).

6. Interpretation, limitations, and legacy

The published analysis emphasized a substantial discrepancy between competition performance and much of the prior literature. Previous works such as FlingBot and UnFoldIR reported coverage around 0.8 and 0.85, but they often used multiple actions, filtered out failed grasps, reset between attempts, or evaluated in more controlled, lab-specific setups. By contrast, the ICRA 2024 competition enforced a single grasp + stretch attempt, included grasp failures in the score, and evaluated all teams on the same physical platform. The resulting lower top coverage, approximately 0.60, was അവതരിപ്പated not as regression but as a more realistic and stringent measure of performance (Gusseme et al., 22 Aug 2025).

Several limitations were acknowledged. The dataset was collected on a single dual-UR5e + ZED2i platform, so viewpoint and kinematic biases remain. The single-attempt constraint does not capture the full potential of iterative unfolding strategies. The dataset contains no explicit cloth-category or semantic-keypoint annotations, although teams can add their own. Some training videos are incomplete because of recording issues, though the 176 competition videos are complete, and the core observation data remain intact (Gusseme et al., 22 Aug 2025).

The benchmark nevertheless provides a concrete foundation for subsequent work. Later real-world comparison frameworks such as DRAPER defined standardized metrics such as Normalised Coverage, Normalised Improvement, and Intersection-over-Union, together with shared robot and fabric protocols for flattening and folding (Kadi et al., 2024). This suggests a convergence toward common cloth-evaluation principles: fixed task protocols, standardized sensing and gripper assumptions, and explicit accounting for failures. In parallel, perception-focused work such as CeDiRNet-3DoF and the ViCoS Towel Dataset showed that competition-aligned public datasets can substantially improve comparability in grasp-point localization (Tabernik et al., 2024).

Within the history of cloth benchmarking, the ICRA 2024 Cloth Competition therefore occupies a specific position. It is not a general benchmark for all deformable-manipulation skills, nor a complete folding or dressing suite. Rather, it is a narrowly scoped, physically grounded benchmark for grasp pose selection on hanging cloth for unfolding, backed by a public interaction dataset and live independent evaluation. Its main contribution is methodological: it converts a historically lab-specific subproblem into a reproducible benchmark with shared data, shared hardware, and directly comparable results, and it does so in a form that the published work explicitly proposes as a foundation for future benchmarks and further progress in data-driven robotic cloth manipulation (Gusseme et al., 22 Aug 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ICRA 2024 Cloth Competition.