HICS-SLAM: Collaborative Semantic SLAM in XR
- The paper demonstrates that integrating real-time human semantic interventions into a robotic SLAM framework significantly improves room detection and localization.
- HICS-SLAM couples a robot-side S-Graphs 2.0 system with a Unity-based XR interface, allowing natural hand gestures for direct 3D scene graph manipulation.
- Experimental results indicate marked gains in semantic completeness, with room recall increasing from 0.38 to 0.95 and improved localization accuracy.
HICS-SLAM, “Human Interaction for Collaborative Semantic SLAM using Extended Reality,” is a human-in-the-loop collaborative semantic SLAM framework in which a robot-side semantic SLAM system and a shared XR environment are coupled so that a human operator can directly inspect and manipulate the robot’s hierarchical 3D scene graph in real time. The system is designed for settings in which semantic SLAM degrades under occlusions, incomplete data, or ambiguous geometries, and it specifically targets the injection of high-level semantic concepts—most notably rooms—into the mapping process. In the reported implementation, robot sensing is handled by S-Graphs 2.0, while human interventions are fused back into the graph as high-confidence constraints, improving semantic completeness and, in some cases, localization and geometric accuracy on construction-site datasets (Ribeiro et al., 18 Sep 2025).
1. Problem formulation and conceptual scope
HICS-SLAM is motivated by a limitation of conventional semantic SLAM: although such systems enrich maps with structural and semantic information, they remain fragile when the environment is only partially observed or structurally ambiguous. The paper identifies several failure modes: occlusions and incomplete data, ambiguous geometry, missing high-level semantics, and error propagation in semantic SLAM. In the stated use cases—especially construction sites and similarly cluttered or unfinished indoor environments—autonomous pipelines may recover planes and walls but still fail to instantiate higher-level entities such as rooms (Ribeiro et al., 18 Sep 2025).
The framework therefore treats the human operator as a source of semantic knowledge that is difficult to extract robustly from sensor data alone. The intended contribution is not low-level manual remapping. Instead, the human supplies high-level spatial interpretation, such as recognizing that a set of partially observed planes defines a room, or reinterpreting a space that an automated pipeline classified as a corridor. The paper explicitly positions this as collaborative semantic SLAM: the robot maps and maintains the scene graph, while the operator augments the map where robot perception is incomplete or ambiguous (Ribeiro et al., 18 Sep 2025).
Extended Reality is introduced as the collaboration medium because it allows the operator to visualize the robot’s 3D scene graph directly in space and to interact with it through natural hand gestures rather than through a conventional 2D parameter-tuning interface. A common misconception would be to treat HICS-SLAM as a visualization layer on top of SLAM. The system is more specific than that: user actions are integrated into the SLAM back end and become graph constraints that affect optimization and the resulting semantic map (Ribeiro et al., 18 Sep 2025).
2. System architecture and scene-graph representation
The robot-side backbone in the reported system is S-Graphs 2.0, described as a metric-relational semantic SLAM system. Robot sensing is performed with 3D LiDAR on a legged platform, using either a Velodyne VLP-16 or an Ouster OS-1 64. The backend organizes the environment as a hierarchical graph containing floors, rooms, walls and planes, and keyframes. The XR application reconstructs this hierarchy in a shared virtual world so that the operator sees a manipulable 3D representation of the robot’s current map (Ribeiro et al., 18 Sep 2025).
The resulting architecture is bidirectional. On the robot side, LiDAR data are processed into geometric and semantic structures by S-Graphs 2.0. On the human side, a Unity-based XR application running on Microsoft HoloLens 2 renders the scene graph and exposes interaction primitives. When the operator adds a room by selecting planes, that room definition is transmitted through ROS 2 back to the SLAM backend, where it is inserted into the hierarchical factor graph and optimized. The updated graph is then propagated back to the XR environment (Ribeiro et al., 18 Sep 2025).
This loop is described as online, real-time, and incremental. The human interventions are confirmation-driven rather than continuously imposed, so the practical interaction mode is asynchronous but online: the robot continues mapping while the operator injects semantic corrections when needed, and those corrections are incorporated incrementally into the current graph state (Ribeiro et al., 18 Sep 2025).
3. XR interaction model and human contribution
The XR interface supports both spatial interactions and semantic interactions. Spatial interactions allow the user to inspect the current map from convenient viewpoints: Zoom in/out is performed with a two-handed pinch gesture, Move with a left-hand pinch gesture, Rotate with a grab gesture, and Select / pinch with a right-hand pinch gesture. The virtual map can therefore be repositioned and examined independent of the physical environment, which is important for semantic inspection of the robot’s internal world model (Ribeiro et al., 18 Sep 2025).
Semantic interaction is centered on room creation. The operator selects four planes that should define a room, confirms the selection, and sends the room definition to the backend. The paper repeatedly emphasizes this operation because it is the implemented semantic capability evaluated experimentally. The human is thus not editing arbitrary graph attributes; the intervention is structured as the addition of a room-level semantic entity inferred from already detected geometric primitives (Ribeiro et al., 18 Sep 2025).
The current implementation is deliberately limited in scope. It supports room semantic integration, not full semantic graph editing. The paper explicitly states that future work will extend the interface toward editing and merging planes, correcting plane identifiers, and resolving duplicate structures. It also claims scalable collaboration in principle, but it does not present an explicit multi-user XR protocol or experimental validation of simultaneous multi-user interaction. In the reported implementation and experiments, the interaction loop is built around a single human operator using HoloLens 2 (Ribeiro et al., 18 Sep 2025).
4. Graph-based semantic fusion and optimization
The core technical contribution of HICS-SLAM is a graph-based semantic fusion mechanism that translates a human semantic intervention into a constraint inside the hierarchical factor graph. When the operator creates a room from four selected planes, the backend inserts a new edge associated with that room and assigns it higher precision weights in the information matrix, reflecting confidence in the human input (Ribeiro et al., 18 Sep 2025).
The paper formalizes the room consistency constraint through a residual tying the room center to the geometric center implied by the selected planes: with
where the paper explains as the distance component from the plane equation of plane , as the unit normal vector of plane pointing inward to the room, and and as opposing plane pairs defining the room’s rectangular horizontal bounds (Ribeiro et al., 18 Sep 2025).
Operationally, this means that the human does not merely attach a label to geometry already in the graph. The added room introduces a structural relation that the optimizer tries to satisfy by reducing the distance between the room-center variable and the center implied by the enclosing planes. The paper further notes that the new rooms are transmitted to S-Graphs without altering the original wall geometry. The primary effect is therefore semantic completion plus a high-confidence geometric consistency constraint at the room level, rather than low-level surface editing (Ribeiro et al., 18 Sep 2025).
A second misconception addressed by the paper follows from this formulation. HICS-SLAM is not presented as a heavily probabilistic semantic fusion model with explicit belief propagation over class labels. The mathematical formalization is light: the main explicit relation is the room-center residual above, together with the statement that the corresponding graph edge is given higher precision in the information matrix. The emphasis is system integration and graph-level intervention rather than a fully elaborated probabilistic semantics model (Ribeiro et al., 18 Sep 2025).
5. Experimental evaluation and reported performance
The experimental setup uses a Unity XR interface, Microsoft HoloLens 2, ROS 2, S-Graphs 2.0, and a legged robot equipped with either a Velodyne VLP-16 or an Ouster OS-1 64 3D LiDAR. Evaluation includes both real-world and simulated datasets. The real-world sets are C1F1, C1F2, C2F0, C2F1, C2F2, C3F1, C3F2, corresponding to floors of compact or larger construction/residential projects. The simulated or mesh-derived sets are SC1F1, SC1F2, SE1, SE2, SE3 (Ribeiro et al., 18 Sep 2025).
The paper evaluates three dimensions: point cloud RMSE for geometric quality on real datasets, Absolute Trajectory Error (ATE) on simulated datasets, and room detection performance using Precision, Recall, and F1-Score against manually established room annotations. The strongest improvements are semantic rather than geometric (Ribeiro et al., 18 Sep 2025).
| Metric | S-Graphs 2.0 | HICS-SLAM |
|---|---|---|
| Average point cloud RMSE | 19.22 cm | 19.15 cm |
| Average ATE | 2.99 | 2.31 |
| Average Precision | ||
| Average Recall | 0 | 1 |
| Average F1-score | 2 | 3 |
On real-world geometry, the gains are modest. The paper reports that HICS-SLAM improves 6 out of 7 real scenarios, with gains up to 1.23% in C1F1 and 1.21% in C2F2, while one case slightly worsens, C2F0: 13.17 → 13.39. The average RMSE changes from 19.22 cm to 19.15 cm, indicating that the method largely preserves baseline geometric quality while adding semantic structure (Ribeiro et al., 18 Sep 2025).
On simulated localization, the average ATE changes from 2.99 to 2.31, reported as an average improvement of 22.7%. The most pronounced reduction is in SE1: 5.07 → 1.15, described as a 77.3% reduction. Other cases show smaller improvements or no change, for example SC1F2 remains 8.29, and SE2 slightly worsens from 1.35 → 1.67. The effect on localization is therefore positive on average but not uniform across all scenarios (Ribeiro et al., 18 Sep 2025).
The dominant result is room detection. Average Recall rises from 0.38 to 0.95, and average F1-score from 0.49 to 0.97. Representative per-dataset values include C2F1 F1: 0.00 → 1.00, C3F1 F1: 0.00 → 0.95, and C3F2 F1: 0.60 → 0.96. Qualitatively, the paper highlights cases in which S-Graphs classified spaces as corridors or only partially detected rooms, while the human operator identified distinct enclosed rooms from the same geometry. These results support the paper’s central claim that the human contribution is primarily a better semantic interpretation of already available geometric evidence (Ribeiro et al., 18 Sep 2025).
6. Relation to prior human-in-the-loop SLAM, limitations, and future directions
HICS-SLAM belongs to a broader human-in-the-loop SLAM lineage, but its emphasis is distinct. Earlier work such as HitL-SLAM incorporated sparse human geometric corrections into pose-graph SLAM by turning approximate human input into correction factors and jointly optimizing the resulting graph (Nashed et al., 2017). HICS-SLAM instead exposes a hierarchical 3D scene graph in XR and uses the operator to add room-level semantic structure to a semantic SLAM backend (Ribeiro et al., 18 Sep 2025). This suggests a shift from human assistance as geometric correction toward human assistance as semantic augmentation.
That distinction also clarifies scope. HICS-SLAM is not a fully autonomous semantic SLAM system, and it is not currently a general-purpose semantic map editor. Its effectiveness depends on operator input quality: if the wrong planes are selected, errors can propagate into the graph. The paper also notes operator variability, a short calibration process for new users, and limited battery life of XR devices such as HoloLens 2. The present implementation only supports room semantic integration, not richer operations such as plane merging or duplicate resolution, and the room model assumes four planes with rectangular horizontal bounds, which does not capture all architectural layouts (Ribeiro et al., 18 Sep 2025).
The future-work agenda follows directly from those limitations. The authors propose extending the system with editing and merging planes, correcting plane identifiers, and resolving duplicate structures, as well as conducting comprehensive user studies and pursuing broader deployment in diverse scenarios. Within the reported formulation, however, the main contribution is already clear: HICS-SLAM shows that a shared XR environment can be used to inject high-confidence room semantics into a hierarchical semantic SLAM graph online, yielding large gains in semantic completeness while preserving, and sometimes slightly improving, the underlying geometric and localization performance (Ribeiro et al., 18 Sep 2025).