CU-Multi: Multi-Robot Dataset
- CU-Multi is a multi-robot dataset designed for evaluating collaborative perception tasks like SLAM and inter-robot data association through independently recorded runs.
- It features synchronized multi-modal sensors and controlled trajectory overlaps to stress both high- and sparse-overlap regimes in outdoor campus environments.
- The dataset provides comprehensive ground-truth, semantic LiDAR annotations, and flexible file organization to support geometry-aware and semantics-aware pipelines.
CU-Multi is primarily a multi-robot dataset introduced in 2025 for evaluating multi-robot data association, collaborative SLAM, map merging, inter-robot loop closure detection, and related collaborative perception tasks. It was collected over multiple days at two large outdoor sites on the University of Colorado Boulder campus and comprises four synchronized runs per environment with aligned start times, controlled trajectory overlap, synchronized multi-modal sensing, semantically annotated LiDAR, and geospatially aligned trajectory references. In the recent literature, the same label also appears in unrelated subfields, including magnetic multilayers and heavy-ion collisions, so its meaning is context dependent (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
1. Definition and research context
CU-Multi was created to address a methodological gap in evaluation for multi-robot systems. The central problem is multi-robot data association: aligning independently collected perception data across space and time when robots observe partially overlapping regions from different poses, at different times, and under different scene conditions. The 2025 dataset papers argue that common practice—splitting a single-robot trajectory into multiple segments—often reuses identical or highly similar viewpoints and lighting, sometimes even the exact same frames, thereby obscuring failure modes and inflating apparent performance in collaborative SLAM and loop-closure pipelines (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
The dataset is therefore structured around independently recorded runs rather than artificial splits. In one description, it is presented as “a dataset for multi-robot data association”; in a later description, it is framed more broadly as “a dataset for multi-robot collaborative perception.” Both descriptions emphasize reproducibility, overlap-aware benchmarking, and realistic pose-dependent variation induced by viewpoint changes, occluders, lighting, and outdoor campus dynamics (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
At dataset scale, CU-Multi contains eight trajectories in total, four per environment, with a combined length of 16.7 km. The longest single trajectory is 4.01 km. One paper names the environments main_campus and kittredge_loop; another refers to env1 (“Main Campus”) and env2 (“Kittredge Loop”) (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
2. Collection design and trajectory structure
CU-Multi uses a single robotic platform to emulate a four-robot team by recording four synchronized runs per environment on different days. The runs were designed to preserve aligned start times and structured overlap relations. A physical wheel-chock procedure was used to align initial pose across runs, and all trajectories terminate near a common rendezvous region. One description states that runs end within about 4 meters of a common location; another reports that they end within 5 m of each other (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
The overlap design is explicit. In each environment, robot1 and robot2 have significant path overlap with different viewpoints; robot3 extends coverage beyond robot1 and robot2 and includes partial overlaps; robot4 is a superset trajectory covering most of the area. The formal trajectory relation is reported as
This structure supports evaluation across both high-overlap and sparse-overlap regimes, rather than a single difficulty level (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
The environment descriptions are also task-relevant. main_campus is a central academic area with approximately 7.4 km of traversed path. kittredge_loop is a large area south of main campus with more open space and varied terrain. Temporal diversity is introduced by multi-day collection, which naturally changes lighting and campus activity. The papers do not enumerate specific weather conditions or dynamic-object distributions, although they note that such factors can be quantified directly from the raw data when needed (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
A central implication of this design is that CU-Multi is not simply a logging corpus; it is an experimental apparatus for overlap-aware benchmarking. High-overlap pairs such as (robot1, robot2) probe viewpoint-invariant data association and map merging, whereas lower-overlap pairs such as (robot1, robot4) stress sparsity, ambiguity, and delayed loop closure (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
3. Sensors, synchronization, and geometric reference structure
The platform is an AgileX Hunter SE with a custom electronics housing and an Intel NUC i7 main computer with 16 GB RAM. The 2025 data-association paper reports the following sensor suite: an Ouster OS1-64 LiDAR with 64 beams, 200 m maximum range, resolution, vertical field of view, and 20 Hz operation; an Intel RealSense D455 RGB-D camera with global-shutter RGB, resolution, field of view, and 30 Hz operation; a MicroStrain 3DM-GQ7 IMU with accelerometer range, 300 dps gyro range, and 400 Hz output; and a MicroStrain 3DM-RTK modem with u-blox ANN-MB-00 antennas providing centimeter-level RTK measurements at 2 Hz and heading accuracy of approximately from modem specifications (Albin et al., 23 May 2025).
Time alignment is handled with Precision Time Protocol (PTP), which is used to align LiDAR, IMU, and compute timestamps. The later collaborative-perception paper also states that LiDAR timestamps are hardware-synchronized to the onboard computer via PTP. This matters because multi-robot evaluation depends not only on geometric overlap but also on timestamp consistency across modalities and runs (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
The calibration package is unusually complete. Reported files include base_to_lidar.txt, base_to_camera.txt, base_to_imu.txt, base_to_rear_left_antenna.txt, and base_to_rear_right_antenna.txt, along with camera intrinsics, meshes, and a URDF in calib/. One release provides OpenStreetMap XML files per environment and sample projection code; another states that ground-truth poses are available both in a local run-centered frame in ROS2 bags and in global UTM coordinates in CSV format with per-line structure
These details make the dataset directly usable for geometry-aware and semantics-aware pipelines without requiring re-estimation of sensor extrinsics (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
4. Ground truth, semantic annotations, and file organization
CU-Multi provides semantically annotated LiDAR and geospatially aligned trajectory references. In the data-association paper, ground-truth odometry is described as tightly-coupled LIO-SAM with RTK GPS, producing consistent odometry and geospatial alignment across runs. In the later collaborative-perception paper, the trajectory reference is refined further through a GTSAM factor graph fusing RTK GPS measurements from both GNSS receivers, LiDAR-inertial odometry from LIOSAM and CT-ICP at 20 Hz, IMU measurements, elevation priors from a 1 m-resolution USGS digital elevation model, and a wheel-chock prior at the known start location. The resulting poses are reported at each LiDAR scan timestamp and aligned to UTM coordinates (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
Semantic LiDAR annotations are generated by a two-stage automated pipeline. Stage 1 uses zero-shot scan-level inference with CENet, a range image-based LiDAR semantic segmentation model, using the SemanticKITTI taxonomy. Stage 2 filters or refines labels using OSM-based geospatial alignment to remove inconsistent labels and enforce label consistency at map scale. Labels are provided per LiDAR scan, and the papers explicitly note that class distribution counts and measured annotation accuracy such as mIoU are not reported; these are intended to be computed by downstream users if required (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
The directory structure is standardized. One release gives the top-level layout as calib/, main_campus/, and kittredge_loop/, with each environment containing robot1/ through robot4/. Each robot directory contains poses.txt, pose_timestamps.txt, camera/data/frame_<number>.bin with timestamps.txt, lidar/data/scan_<number>.bin, lidar/labels/scan_label_<number>.bin, imu/data.txt, and gps/data.txt. The later release instead emphasizes ROS2 .db3 bags for RGB-D, LiDAR, IMU, GNSS, and ground truth, plus per-robot CSV files with UTM poses, together with conversion utilities to ROS1 bag and ROS2 .mcap (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
The overlap structure is also partially quantified. The data-association paper reports pairwise “qualitative overlap metrics” in a dedicated table. For example, in Main Campus, target robot1 has overlap magnitudes 1136.41 with source robot1, 788.43 with robot2, 538.11 with robot3, and 246.89 with robot4; in Kittredge Loop, target robot1 has corresponding values 1295.96, 595.97, 371.70, and 523.19. The authors describe these as indicative magnitudes of shared path or co-visibility and recommend reporting explicit overlap measures rather than relying on arbitrary trajectory splits (Albin et al., 23 May 2025).
5. Benchmarking tasks, metrics, baselines, and limitations
CU-Multi is intended for inter-robot loop closure detection, multi-robot data association, collaborative SLAM, map merging, LiDAR place recognition, and semantic mapping. The papers recommend evaluation regimes that reflect the designed overlap structure: (robot1, robot2) for high-overlap and viewpoint-diverse tests, (robot3, robot4) for sparser overlap, and rendezvous regions as high-confidence ground-truth closures (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
For loop closure detection, the recommended metrics are Precision, Recall, and :
For multi-robot matching, the recommended binary assignment formulation is
0
subject to 1, 2, and 3, with solution by Hungarian algorithm or min-cost flow. For collaborative SLAM and map merging, the pose-graph objective is given as
4
with 5 denoting relative-pose residuals and 6 a robust loss such as Huber. Trajectory quality is reported with ATE and RPE, and semantic performance with mIoU (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
The later paper adds baseline methods. It reports a LiDAR place-recognition baseline using ScanContext, with a loop closure counted as correct if ground-truth separation is 7 m, and a distributed C-SLAM baseline using DiSCo-SLAM extended to four robots. To avoid immediate inter-robot closures, playback was offset by 200 s. Reported DiSCo-SLAM numbers on CU-Multi are:
- Main Campus: ATE (m) 8 5.58, 8.52, 14.30, 22.85 for 9–0; RPE (m) 1 0.11, 0.11, 0.49, 0.31.
- Kittredge Loop: ATE (m) 2 7.68, 5.01, 14.42, 18.85 for 3–4; RPE (m) 5 0.08, 0.16, 0.35, 0.45.
These values are explicitly reported only for runs that did not diverge due to erroneous closures (Albin et al., 23 Sep 2025).
The limitations are equally explicit. CU-Multi uses a single platform across multiple sessions rather than truly simultaneous multi-robot acquisition. The release contains two outdoor campus environments only, with no indoor sequences or adverse-weather runs. Semantic labels are automated through zero-shot CENet plus OSM refinement and may include residual errors or class biases. GNSS quality degrades near dense buildings, especially in elevation, although the factor-graph fusion is intended to mitigate this. The later paper also states that it does not prescribe official train/validation/test splits, although it proposes a reproducible recommendation using high-overlap trajectories for training and low-overlap trajectories for testing (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).
6. Other uses of the label in the literature
The label “CU-Multi” is not unique to robotics. It appears in several unrelated technical contexts, and disambiguation is often necessary.
| Usage | Domain | Meaning |
|---|---|---|
| CU-Multi | Multi-robot perception | Dataset for multi-robot data association and collaborative perception |
| CU-Multi | Magnetic multilayers | Electrodeposited Co/Cu multilayers studied for AMR-to-GMR evolution |
| CU-Multi | Heavy-ion collisions | Fraction of Cu+Cu participants undergoing multiple inelastic collisions |
| Cu-Multi | Bioactive thin films | Ti/Cu/Ti multilayer thin-film system patterned by femtosecond laser |
In electrodeposited Co/Cu multilayers, the term refers to a sample family with Cu spacer thickness varied from 0.5 to 4.5 nm. That study reports an AMR-to-GMR transition near 6 nm, a monotonic increase of saturation GMR to about 5–6% at 7–4.0 nm, and the absence of oscillatory interlayer exchange coupling. The mechanistic explanation is microstructural: pinholes, roughness, thickness fluctuations, and intermixing suppress RKKY-like oscillatory coupling and favor ferromagnetic coupling at small spacer thicknesses (Bakonyi et al., 2016).
In STAR heavy-ion analyses of strangeness enhancement, “CU-Multi” denotes the Glauber-model-derived fraction of participating nucleons in Cu+Cu collisions at 8 GeV that undergo multiple inelastic collisions, 9, where 0. It enters the core–corona parameterization
1
used to describe system-dependent strangeness enhancement. The study reports that at matched 2, strange-hadron production per participant is higher in Cu+Cu than in Au+Au, and attributes this to the larger fraction of multiple-collision participants in Cu+Cu (Collaboration et al., 2011).
A case-variant, “Cu-Multi,” is used for Ti/Cu/Ti multilayer thin films on Si patterned by ultrafast laser processing. In that context, the term denotes a 300 nm multilayer with a 10 nm Cu subsurface layer and a 10 nm Ti top layer. The reported 3 mm mesh patterns exhibit less well-defined LSFL than Ti/Zr/Ti, local HSFL with 4 nm, Cu depletion in irradiated zones, increased oxygen content along laser-written lines, and a slight 5 decrease in MRC-5 viability relative to control while remaining non-cytotoxic overall (Petrović et al., 2023).
These unrelated usages do not share a common technical meaning beyond the label itself. In current arXiv usage, however, the term is most prominently associated with the University of Colorado multi-robot dataset introduced in 2025 (Albin et al., 23 May 2025, Albin et al., 23 Sep 2025).