---
title: GrandTour Dataset for Legged Robotics
url: https://www.emergentmind.com/topics/grandtour-dataset
type: topic
---

# GrandTour Dataset for Legged Robotics

GrandTour is a multi-modal legged-robotics dataset collected on an ANYbotics ANYmal-D quadruped equipped with the Boxi multi-modal sensor payload, designed for research on SLAM, high-precision state estimation, multi-modal perception, and navigation in complex outdoor and indoor environments [2602.18164]. The dataset spans 49 missions, exceeds 5 hours of recorded data and 10 km of traversed distance, and combines exteroceptive sensing, proprioception, and high-precision reference trajectories under conditions such as snow glare, dense foliage, debris, dust, dark interiors, dynamic urban scenes, and rapid legged-motion disturbances [2602.18164]. Subsequent studies use GrandTour not only as a broad perception benchmark but also as a controlled testbed for sensor-configuration studies and proprioceptive-only state-estimation benchmarks, which together make the dataset notable for both modality breadth and evaluation diversity [2606.19067].

## 1. Scope, platform, and mission coverage

GrandTour comprises 49 missions across four broad environment categories: Alpine, Forest, Demolished or Industrial Buildings, and Urban and Indoor [2602.18164]. The Alpine subset includes Jungfraujoch and Eiger station hikes, labeled with prefixes such as SNOW-*, EIG-*, and PIL-*; the Forest subset includes Forêt de Montmorency, Känzeli, Trimstein, and Albträsgarten, labeled CYN-*, KÄB-*, TRIM-*, and ALB-*; the industrial and demolished-building subset includes search-and-rescue and construction-site sequences such as ARC-*, SPX-*, and CON-*; and the Urban and Indoor subset includes ETH campus, historic quads, and warehouse environments such as ETH-*, GRI-*, LEICA-*, HAUS, HÖB, and LEE [2602.18164].

The reported challenges vary by category. Alpine missions include snow glare, featureless expanses, and high-altitude thin air. Forest missions include dense foliage, uneven terrain, and mixed lighting. Demolished and industrial missions include tight corridors, debris, dust or smoke, and dark interiors. Urban and indoor missions include cars, pedestrians, loop closures, and illumination changes [2602.18164]. In the systematic multimodal SLAM evaluation on GrandTour, additional terrain descriptions include pavement, grass, gravel, mud, snow, stairs, wet surfaces, and tunnel transitions, together with legged-motion disturbances such as foot-impact shocks at approximately 20 Hz, high-frequency mechanical vibrations up to 100 Hz, and rapid yaw rotations above \(150^\circ/\mathrm{s}\) during turns and reverse-pass loops [2606.19067].

The platform is an ANYmal D quadruped, with sensors mounted both on the Boxi payload and on the robot body itself [2602.18164]. In one documented Boxi arrangement used for multimodal SLAM evaluation, Sevensense stereo cameras are mounted at approximately 0.8 m above ground with a 12 cm baseline, a ZED2i unit is installed immediately above them with the same yaw and pitch orientation, a Livox Mid-360 is mounted above both cameras, and dual IMUs are co-located within the Boxi core in a body frame oriented \(x\)-forward, \(y\)-left, \(z\)-up [2606.19067]. This suggests that GrandTour is not merely a collection of trajectories but a platform-specific dataset in which sensor placement and embodiment are part of the experimental object.

## 2. Sensor modalities and instrumentation

GrandTour’s sensor suite includes spinning 3D LiDARs, RGB cameras, depth cameras, multiple IMUs, onboard proprioceptive signals, and global-positioning and reference systems [2602.18164]. The spinning 3D LiDARs are the Livox Mid-360, the Hesai XT-32, and the Velodyne VLP-16. Their reported rates are 10 Hz each, with respective vertical fields of view of approximately \(77.2^\circ\), \(31^\circ\), and \(30^\circ\), horizontal fields of view of \(360^\circ\), and stated range and accuracy values of \(0.1\text{–}40\) m with \(\pm 0.02\) m, \(0.05\text{–}120\) m with \(\pm 0.01\) m, and \(0.5\text{–}100\) m with \(\pm 0.03\) m [2602.18164].

The RGB camera suite includes five Sevensense CoreResearch cameras, three TierIV C1 HDR cameras, and a Stereolabs ZED2i stereo camera [2602.18164]. The CoreResearch cameras are global-shutter, \(1440\times1080\) px, 10 Hz, with \(126^\circ\times92.4^\circ\) FoV and HDR. The TierIV C1 HDR units are rolling-shutter, \(1920\times1280\) px, 30 Hz, with \(120^\circ\times80^\circ\) FoV and 120 dB dynamic range. The ZED2i stereo camera is rolling-shutter, \(1920\times1080\) px, 15 Hz, with \(110^\circ\times70^\circ\) FoV [2602.18164]. Depth sensing is provided by six Intel RealSense D435i units mounted on ANYmal and by ZED2i depth and confidence images at 15 Hz [2602.18164].

The proprioceptive stack includes several IMUs with different grades and placements, among them the Honeywell HG4930 IMU in the CPT7 at 100 Hz, the Safran STIM320 at 500 Hz, the TDK ICM40609 in the Livox at 200 Hz, the Bosch BMI085 in the CoreResearch at 200 Hz, the ADIS16475-2 at 200 Hz, the ZED2i IMU at 45 Hz, ANYmal’s onboard IMU at 400 Hz, and 12 joint encoders providing position, velocity, and foot-contact signals [2602.18164]. A later controlled evaluation isolates two IMU tiers in particular: the Analog Devices ADIS15475-2, described as industrial-grade, with 200 Hz output, gyro noise density of approximately \(0.02^\circ/\sqrt{\mathrm{Hz}}\), and accelerometer noise density of approximately \(60\,\mu g/\sqrt{\mathrm{Hz}}\); and the Honeywell HG4930, described as tactical-grade, with 100 Hz output, gyro noise density of approximately \(0.005^\circ/\sqrt{\mathrm{Hz}}\), and accelerometer noise density of approximately \(30\,\mu g/\sqrt{\mathrm{Hz}}\) [2606.19067].

A common misconception is that GrandTour is primarily an exteroceptive dataset. The released descriptions do not support that view: the platform combines LiDAR, multiple RGB and depth cameras, multiple IMUs, joint encoders, and contact signals, and one published benchmark explicitly evaluates proprioceptive-only state estimation on the CYN-1 sequence [2605.11674].

## 3. Synchronization, calibration, and reference trajectories

All GrandTour sensors are described as rigidly calibrated and time-synchronized to a common clock [2602.18164]. In the full dataset description, the NovAtel CPT7 GNSS receiver acts as PTP grandmaster, and Jetson AGX Orin, Intel NUC, and Raspberry Pi networks are PTP-synchronized to the CPT7 under IEEE 1588 v2 PTP, with stated sub-\(\mu\)s to ns accuracy [2602.18164]. Hardware triggers are used for IMUs and cameras, with exposure-centered timestamps, while USB devices use software timestamps with stated jitter of at most 1 ms [2602.18164]. In the multimodal SLAM evaluation release, all sensors are hardware-time-stamped and synchronized via the Boxi payload’s FPGA time bus, and timestamps are emitted in ROS-compatible UNIX nanosecond clock [2606.19067].

The synchronization model in the dataset description is written as
\[
t_\mathrm{sync}^A = t_\mathrm{host}^A + \delta_A,\quad
t_\mathrm{sync}^B = t_\mathrm{host}^B + \delta_B,
\]
with nearest-neighbor merging under
\[
|t_\mathrm{sync}^A - t_\mathrm{sync}^B| \le t_\mathrm{tol}.
\]
Rolling-shutter cameras are timestamped at the midpoint of exposure, while all global-shutter cameras trigger simultaneously under PTP [2602.18164].

Calibration procedures are reported in detail for the Boxi-based SLAM evaluation setup. Camera intrinsics are modeled via the Kannala–Brandt projection. Extrinsics are obtained through target-based checkerboard calibration followed by hand–eye refinement in Kalibr style. LiDAR–camera calibration uses depth-to-image alignment via iterative closest point on static scene captures, and IMU–camera alignment uses static orientation holds and cross-covariance minimization [2606.19067]. These details are relevant because later benchmark findings explicitly tie performance differences to sensor modality, shutter type, and inertial integration choices.

GrandTour provides high-precision reference trajectories from satellite-based RTK-GNSS and a Leica Geosystems total station [2602.18164]. The NovAtel SPAN CPT7 dual-antenna RTK GNSS with PPP corrections operates at 20 Hz in real time and is post-processed with Inertial Explorer. Reported short-outage performance includes horizontal error of at most 0.02 m RMS and attitude error of at most \(0.01^\circ\) for outages up to 10 s [2602.18164]. The Leica MS60 total station with GRZ101 mini-prism via AP20 auto-pole provides 20 Hz 3D positions with stated accuracy of \(\pm 1.5\) mm \((3\sigma)\), with AP20 timestamps synchronized to PTP within 1 ms [2602.18164].

The full dataset also describes a holistic fusion factor graph for reference-trajectory estimation:
\[
X^*=\arg\max_X\,p(X\mid Z)=\arg\max_X\,p(Z\mid X)\,p(X),
\]
where \(X=\{\,^I X_n,\,^G X_n,\,^R X_n\}\) are robot poses in inertial, GNSS, and reference frames, with factors for IMU preintegration, GNSS unary, TPS unary, and gravity priors [2602.18164]. In a later SLAM-oriented release, indoor ground truth is given by Vicon motion capture at 100 Hz, with only position used because roll and pitch are noisy, while outdoor ground truth uses RTK-GPS combined with LiDAR-SLAM corrected trajectories with at most 0.05 m drift per 100 m [2606.19067]. This indicates that “ground truth” in GrandTour is modality- and release-dependent rather than a single uniform pipeline.

## 4. Data products, file organization, and access paths

GrandTour is distributed in several formats. The full release is available at the project website, on HuggingFace in a ROS-independent format, and in ROS formats [2602.18164]. The HuggingFace version uses per-topic Zarr arrays with timestamps, intrinsics, and extrinsics stored under attributes and metadata, together with compressed images in JPEG or PNG [2602.18164]. ROS 1 bags are LZ4-compressed, with 34 bags per mission, one per sensor or derived stream, a `tf_static` bag with extrinsics, and each bag namespace under `/boxi` or `/anymal`; a provided Python script converts the data to ROS 2 `.mcap` [2602.18164]. Download access is also exposed through the Kleinkram Data Management CLI and web interface by mission UUID and topic pattern [2602.18164].

A later public release associated with the multimodal SLAM study provides a per-mission directory layout of the form
`/MXX/`
with `camera_left/`, `camera_right/`, `camera_zed_left/`, `camera_zed_right/`, `depth_zed/`, `imu_adis.csv`, `imu_honeywell.csv`, `lidar/`, `gt_trajectory.csv`, `calib/`, and `timestamps.txt` [2606.19067]. The image formats are specified as 16-bit mono PNG for the Sevensense pair, JPEG or PNG RGB for the ZED images, and 16-bit PNG depth in millimeters for `depth_zed/`; LiDAR data are stored as binary PointCloud2 `.bin` with `x,y,z,intensity` [2606.19067]. Each file’s first column is the hardware timestamp in nanoseconds, and a master index file lists all sensors for a given mission in time-sorted order [2606.19067].

The same release provides utilities for converting ROS bags to KITTI format and then to the EVO evaluation toolkit, ROS launch files for replaying each mission, and calibration loader nodes for ORB-SLAM3, RTAB-Map, FAST-LIVO2, and DPV-SLAM [2606.19067]. The quick-start procedure consists of cloning the repository, downloading raw data bundles via `download.sh`, replaying a mission with `roslaunch grandtour replay_m10.launch`, launching a SLAM stack against replay topics such as `/camera/left/image_raw`, `/camera/right/image_raw`, `/imu/data`, and `/livox/point_cloud`, and then using `evaluate_ate_rpe.py` against `gt_trajectory.csv` [2606.19067].

Published materials also document mission-specific examples. A seven-sequence subset totaling approximately 2731 s, or about 45 min, includes M10 (Snowy alpine, 236 s), M13 (Urban with pedestrians/cars, 455 s), M19 (Mountain trail + reverse loop, 409 s), M24 (Industrial muddy site, 290 s), M34 (Outdoor \(\rightarrow\) underground transition, 529 s), M42 (Outdoor \(\rightarrow\) indoor with dynamic obstacles, 310 s), and M44 (Industrial railway long pan, 502 s) [2606.19067].

## 5. Evaluation protocols and benchmark metrics

GrandTour is associated with multiple benchmark protocols rather than a single canonical evaluation. In the original large-scale benchmark, 52 open-source VO, VIO, LO, and LIO methods were evaluated on six representative sequences: SPX-2, SNOW-2, EIG-1, CON-4, ARC-2, and ARC-7 [2602.18164]. Two primary metrics were used. Absolute Trajectory Error was defined as
\[
\mathrm{ATE}=\sqrt{\frac{1}{N}\sum_{i=1}^N \left\| \mathbf p_i^\mathrm{est}-\mathbf p_i^\mathrm{gt}\right\|^2},
\]
and Relative Translation Error over path length \(\Delta=0.5\) m was defined as
\[
\mathrm{RTE}(\Delta)=\sqrt{\frac{1}{M}\sum_{j=1}^M
\left\|(\mathbf p_{j+\Delta}^\mathrm{est}-\mathbf p_j^\mathrm{est})-
(\mathbf p_{j+\Delta}^\mathrm{gt}-\mathbf p_j^\mathrm{gt})\right\|^2 }.
\]
Trajectories were aligned via Umeyama’s 7-DoF least-squares, and timestamp association used a nearest-neighbor threshold \(t_{\max\_diff}=\tfrac12\mathrm{median}(\Delta t)\) [2602.18164].

The later multimodal SLAM evaluation on ANYmal D uses ATE and Relative Pose Error in a different but related form. After rigid alignment in \(SE(3)\) using Umeyama, ATE is
\[
\mathrm{ATE}=\sqrt{\frac{1}{N}\sum_{i=1}^{N}\left\|p_i^\mathrm{gt}-T^*\cdot p_i^\mathrm{est}\right\|^2},
\]
and with interval \(\Delta\), RPE is
\[
\mathrm{RPE}=\sqrt{\frac{1}{N-\Delta}\sum_{i=1}^{N-\Delta}
\left\|\mathrm{trans}\!\left((\Delta T_i^\mathrm{gt})^{-1}\Delta T_i^\mathrm{est}\right)\right\|^2 }.
\]
Reference trajectories are stored in `gt_trajectory.csv` as time series of positions \((x,y,z)\) in the Boxi body frame [2606.19067].

A proprioceptive-only benchmark on the GrandTour CYN-1 sequence uses yet another published protocol. CYN-1, also called Grindelwald Canyon, is an outdoor GNSS-enabled mission of approximately 296 m over uneven, rock-strewn canyon terrain with elevation changes, tight turns, and intermittent foot slips [2605.11674]. The benchmark evaluates MUSE, IEKF, and the Invariant Smoother using evo 1.10, with ATE, translational and rotational RPE over \(\Delta=1\) m and \(\Delta=1\) frame, velocity RMSE, and per-update runtime [2605.11674]. The preprocessing pipeline consists of converting ROS bags to `sensor_data.csv` and `groundtruth.csv`, precomputing foot position \(d_\ell(t)\), Jacobian \(J_\ell(q)\), and foot velocity \(v_\ell(t)\) with Pinocchio, aligning timestamps across IMU, encoder, contact, and ground-truth streams by linear interpolation, and exporting estimated trajectories in TUM format for evo evaluation [2605.11674].

These differing evaluation definitions do not contradict one another; they indicate that GrandTour functions as a dataset substrate for several research questions, including global drift, short-horizon consistency, and computational cost.

## 6. Empirical findings, interpretations, and research uses

The broad GrandTour benchmark reports that LiDAR-inertial odometry methods such as Coco-LIC and FAST-LIVO2 outperform pure LiDAR odometry and visual-inertial methods, especially in dynamic and feature-poor scenes [2602.18164]. Multi-LiDAR methods such as CTE-MLO improve robustness in confined spaces, while VIO methods fail under dark or extreme lighting in ARC-7, which highlights the need for re-initialization [2602.18164]. Online extrinsic re-calibration is reported to have mixed benefits, with observability described as motion dependent [2602.18164]. The dataset paper recommends starting with a robust LIO baseline such as FAST-LIVO2 or Coco-LIC, disabling loop closure for pure drift benchmarks, tuning initialization, checking IMU–camera and LiDAR calibration, incorporating time-synchronization verification, and evaluating across all four environment categories [2602.18164].

The sensor-configuration study adds more specific findings about GrandTour’s relevance for legged locomotion. Across visual, visual-inertial, and LiDAR-visual-inertial SLAM methods, stereo configurations consistently outperform monocular and RGB-D modalities, global-shutter cameras significantly mitigate motion-induced tracking failures compared with rolling-shutter cameras, and standard inertial integration can degrade the performance of primarily vision-based frameworks under harsh legged locomotion [2606.19067]. Those conclusions are explicitly linked to the embodiment-induced sensory challenges of quadrupeds, including foot-impact shocks, high-frequency vibrations, and rapid angular rotations [2606.19067].

The proprioceptive-only CYN-1 benchmark further shows how GrandTour can be used without exteroception. On that sequence, ATE is reported as 2.269 m for MUSE, 1.406 m for IEKF, and 1.363 m for the Invariant Smoother with \(WS=3\); velocity RMSE is 0.876, 0.869, and 0.869 m/s, respectively; RPE at \(\Delta=1\) m is 0.0722, 0.0432, and 0.0425 m; and mean per-iteration runtime is \(0.012\pm0.002\) ms for MUSE, \(0.020\pm0.004\) ms for IEKF, and \(0.260\pm0.020\) ms for the Invariant Smoother with \(WS=3\) [2605.11674]. The accompanying discussion attributes lower ATE for IEKF and IS to group-affine contact modeling and notes a latency–accuracy trade-off across filters and smoothers [2605.11674].

The dataset description lists a wide range of intended applications: legged SLAM and odometry, multi-modal representation learning, vision-based locomotion and physical parameter estimation, real-to-sim and domain transfer with neural scene representations, 3D volumetric and mesh mapping, foundational model fine-tuning for embodied navigation, dynamic object segmentation and dynamic SLAM, and end-to-end navigation stack benchmarking [2602.18164]. A plausible implication is that GrandTour is particularly useful where modality interaction, synchronization quality, and embodiment-specific disturbances are first-order variables rather than nuisance factors.

A second common misconception is that GrandTour should be treated as a single benchmark with a fixed sensor stack and fixed metrics. The published material instead describes a larger dataset with multiple access formats and mission categories, a controlled multimodal SLAM subset with its own file structure and ground-truth conventions, and a proprioceptive benchmark centered on CYN-1 [2602.18164]. For research practice, that distinction matters: results obtained on one GrandTour-derived protocol are not automatically interchangeable with results obtained on another, even when they share the same underlying platform and mission family.

Source: https://www.emergentmind.com/topics/grandtour-dataset