Papers
Topics
Authors
Recent
Search
2000 character limit reached

Verti-Arena: Indoor Off-Road Benchmark

Updated 8 July 2026
  • Verti-Arena is a reconfigurable indoor facility designed for off-road autonomy research, offering a standardized benchmark through modular terrain configurations and precise ground-truth measurements.
  • The facility enables remote, reproducible experiments using a web-based interface and a detailed sensing stack including RGB-D cameras, IMUs, and motion-capture systems.
  • Empirical findings show that specialized data-driven models outperform classical kinematic models in predicting vehicle dynamics across diverse off-road conditions.

Verti-Arena is a reconfigurable indoor facility designed specifically for off-road autonomy, introduced to address the lack of a controllable and standardized real-world testbed for systematic data collection and validation in off-road navigation research. It provides a repeatable benchmark environment for reproducible experiments across vertically challenging terrains, precise ground-truth measurement through onboard sensors and a motion-capture system, and a web-based interface for remotely conducting standardized experiments (Chen et al., 11 Aug 2025).

1. Research role and problem setting

Off-road navigation is framed as a core capability for mobile robots operating in environments that are inaccessible or dangerous to humans, including disaster response and planetary exploration. The central limitation identified for the field is not only algorithmic difficulty but also experimental infrastructure: progress is constrained by the absence of a controllable and standardized real-world testbed that supports systematic data collection, validation, and comparative evaluation (Chen et al., 11 Aug 2025).

Within that setting, Verti-Arena is positioned as an indoor benchmark facility rather than merely a robotics course. Its function is to make experiments repeatable across terrain conditions, to support consistent data collection, and to permit comparative evaluation of off-road autonomy algorithms under shared conditions. This suggests a research emphasis on inter-run comparability, cross-model benchmarking, and reproducibility across laboratories rather than solely on one-off demonstrations.

2. Physical configuration and terrain modularity

Verti-Arena spans an 8m×8m8\,\mathrm{m}\times 8\,\mathrm{m} footprint and supports a maximum elevation change of 0.7m0.7\,\mathrm{m}. The floor is composed of a 4×44\times 4 grid of 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m} terrain panels, each of which can be individually swapped or rearranged. Elevation features are constructed from interlocking plywood wedges and foam inserts that lock into dovetail channels on each panel, enabling the formation of hills, cliffs, ravines, and ramps. Slopes and steps can be configured with varying pitch up to 3030^\circ (Chen et al., 11 Aug 2025).

The surface semantics are also modular. Semantic modules include sand beds, stone-dust mats, fine pebbles, grass turf, flagstone tiles, wood planks, foam board, concrete slabs, and artificial trees, mounted on magnetized sub-panels for rapid reconfiguration. To generate a new layout, operators choose which panels carry which semantic cover, insert elevation blocks, and place obstacles such as boulders greater than 0.3m0.3\,\mathrm{m} diameter, wooden barriers, and shrub clusters. The arena supports thousands of possible layouts (Chen et al., 11 Aug 2025).

This combination of topographic variation and semantic variation is central to the facility’s design. Standardization is achieved at the level of panelized assembly, measurement, and protocol, while environmental diversity is achieved through reconfiguration. A plausible implication is that Verti-Arena standardizes how terrain is specified and reproduced, not by collapsing terrain diversity into a single canonical layout.

3. Sensing stack, coordinate frames, and calibration

The experimental platform uses a V4W four-wheeled vehicle instrumented with onboard sensors and external motion capture. The onboard RGB-D sensor is a Microsoft Azure Kinect with color at 1280×720×31280\times 720\times 3 up to 30Hz30\,\mathrm{Hz} and depth at 512×512512\times 512 at 30Hz30\,\mathrm{Hz}; intrinsic parameters 0.7m0.7\,\mathrm{m}0 are stored in each ROS 2 bag. The IMU provides a 3-axis gyroscope and accelerometer at 0.7m0.7\,\mathrm{m}1 with noise density approximately 0.7m0.7\,\mathrm{m}2. Wheel encoders and joint states provide four wheel speeds and steering angles at 0.7m0.7\,\mathrm{m}3. Computation is performed on an NVIDIA Jetson Xavier NX running ROS 2 Foxy. External ground truth is supplied by eight Vicon or OptiTrack cameras surrounding the arena, sampling at 0.7m0.7\,\mathrm{m}4 with sub-millimeter translational accuracy and sub-degree orientation accuracy (Chen et al., 11 Aug 2025).

Component Specification Rate / accuracy
RGB-D camera Microsoft Azure Kinect Color up to 30 Hz; depth 30 Hz
IMU 3-axis gyroscope and accelerometer 100 Hz; 0.7m0.7\,\mathrm{m}5
Wheel encoders & joint states 4 wheel speeds and steering angles 100 Hz
Motion capture 8 Vicon or OptiTrack cameras 100 Hz; sub-mm / sub-degree

The coordinate-frame conventions follow ROS 2 TF practice. The world frame 0.7m0.7\,\mathrm{m}6 has origin at the center of the arena floor, with 0.7m0.7\,\mathrm{m}7 forward, 0.7m0.7\,\mathrm{m}8 left, and 0.7m0.7\,\mathrm{m}9 up. The base link 4×44\times 40 is defined at the geometric center between wheel axles with the same axis orientation. The camera frame 4×44\times 41 uses standard Kinect coordinates, with 4×44\times 42 forward from the camera lens. The motion-capture system publishes TF2 transforms /world → /mocap_base_link and /mocap_base_link → /base_link, the latter obtained through static calibration (Chen et al., 11 Aug 2025).

Calibration is formalized in two places. Camera–IMU extrinsics are calibrated via the Kalibr toolbox. The mocap-to-vehicle transform 4×44\times 43 is obtained via a one-time hand-eye calibration:

4×44\times 44

where 4×44\times 45 is an OptiTrack marker frame, 4×44\times 46 the camera frame, and 4×44\times 47 the base link. This calibration chain is what allows the platform to align onboard estimates with external ground truth in a common reference system.

4. Ground truth, vehicle models, and evaluation formalism

The motion-capture system provides the 4×44\times 48-DoF pose 4×44\times 49 of the vehicle base_link at 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}0, while onboard odometry 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}1 is recorded from wheel encoders and IMU. Vehicle kinodynamics are formalized with the forward model

2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}2

where 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}3 and 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}4, with 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}5 denoting longitudinal velocity and 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}6 steering rate (Chen et al., 11 Aug 2025).

The reported evaluation formalizes multiple error metrics. Pose error at time 2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}7 is

2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}8

Orientation error is

2m×2m2\,\mathrm{m}\times 2\,\mathrm{m}9

Velocity error is

3030^\circ0

In the reported experiments, mean positional RMSE and angular RMSE are plotted over each trajectory segment (Chen et al., 11 Aug 2025).

Benchmarking is organized around five terrain zones: boulder, flagstone, stone dust, grass, and pebble. For each zone, 3030^\circ1 runs are executed over the same start-goal waypoint pair in low-gear mode with locked differentials. Success Rate is defined as

3030^\circ2

Traversal Time for run 3030^\circ3 is

3030^\circ4

and Average Slip Error is

3030^\circ5

where 3030^\circ6 is the slip ratio estimated from wheel encoders versus ground-truth velocity. Model comparison is performed between classical bicycle and Ackermann kinematic models and zone-specific MLPs. Per-zone RMSE reported in Fig. 7 places MLP positional RMSE in the range 3030^\circ7–3030^\circ8 and classical models in the range 3030^\circ9–0.3m0.3\,\mathrm{m}0; the terrain difficulty ranking from highest RMSE to lowest is boulder 0.3m0.3\,\mathrm{m}1 pebble 0.3m0.3\,\mathrm{m}2 stone dust 0.3m0.3\,\mathrm{m}3 grass 0.3m0.3\,\mathrm{m}4 flagstone (Chen et al., 11 Aug 2025).

5. Remote experimentation and reproducible data collection

A web-based remote interface allows research groups worldwide to conduct standardized experiments. The workflow comprises authentication and reservation of an experiment slot through a web portal, upload of autonomy nodes as ROS 2 packages or Docker containers, configuration of terrain layout through presets or custom JSON describing panel arrangement, selection of start/goal coordinates and run count, and choice between teleop and autonomy mode. Experiments can then be launched remotely, monitored through live RGB and depth video feeds together with streaming telemetry including pose, control commands, and error plots. After completion, users download recorded ROS 2 bag files containing all sensors, mocap, and TF2 trees, as well as summary CSV logs of metrics (Chen et al., 11 Aug 2025).

The reproducibility layer is specified in operational detail. All data are stored in ROS 2 bag2 format. Standard topic names include /camera/color/image_raw at 0.3m0.3\,\mathrm{m}5, /camera/depth/image_raw at 0.3m0.3\,\mathrm{m}6, /imu/data at 0.3m0.3\,\mathrm{m}7, and /tf and /tf_static as geometry_msgs/TransformStamped at 0.3m0.3\,\mathrm{m}8. Global clock synchronization is maintained via Chrony to less than 0.3m0.3\,\mathrm{m}9 offset, and motion capture and onboard sensors are triggered by PTP, ensuring sub-millisecond alignment. Recommended operating parameters are low-gear speed 1280×720×31280\times 720\times 30, steering rate limited to 1280×720×31280\times 720\times 31 to avoid drivetrain slip beyond 1280×720×31280\times 720\times 32, and locked differentials for baseline data, with unlocked differentials reserved for advanced traction-control studies (Chen et al., 11 Aug 2025).

These design choices make reproducibility a property of both the environment and the data products. The standardization extends from terrain layout and vehicle state estimation to packaging, synchronization, logging, and post hoc retrieval.

6. Empirical findings and research significance

The reported forward-model evaluation shows that per-zone MLPs achieve mean positional RMSE 1280×720×31280\times 720\times 33 on flagstone and approximately 1280×720×31280\times 720\times 34 on boulders, while classical kinematic models uniformly exceed 1280×720×31280\times 720\times 35 RMSE and fail to predict rollovers and wheel lift-off in 1280×720×31280\times 720\times 36-DoF motions. The results are interpreted as evidence that zone-specialized data-driven models capture subtle terrain–vehicle interactions, including slip bursts on pebbles. The same evaluation also indicates that blended semantics, such as grass over flagstones, introduce mixed dynamic regimes for which training on pure zones only partially generalizes (Chen et al., 11 Aug 2025).

A separate empirical observation concerns repeatability. Remote users reportedly reproduce runs within 1280×720×31280\times 720\times 37 variance in traversal time across weeks of experiments. This suggests that the facility’s claims to controllability are not limited to local operation by the host laboratory but extend to remote use under standardized protocols (Chen et al., 11 Aug 2025).

A plausible misconception is that standardization in off-road robotics necessarily entails simplification into flat or low-variance terrain. Verti-Arena instead combines a fully modular 1280×720×31280\times 720\times 38 arena, ten terrain types, up to 1280×720×31280\times 720\times 39 elevation change, a millimeter-accurate ground-truth pipeline, and open benchmarking scenarios with success, traversal, and slip metrics. In that sense, its significance lies in making multi-terrain off-road autonomy experimentally comparable without removing the vertically structured and semantically heterogeneous conditions that generate the problem in the first place (Chen et al., 11 Aug 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Verti-Arena.