TWC-SLAM: Multimodal Cooperative LiDAR SLAM
- TWC-SLAM is a multi-agent cooperative LiDAR SLAM framework that fuses LiDAR+IMU data with text semantics and WiFi fingerprints to resolve aliasing in similar indoor layouts.
- It employs a multimodal gating strategy where text similarity and WiFi fingerprint matching generate candidate co-locations that are then verified using LiDAR registration via ICP.
- Experiments in repetitive indoor settings demonstrate that TWC-SLAM reduces endpoint errors to 0.16–0.21 m, outperforming single-modality approaches in loop closure and map alignment.
Searching arXiv for the TWC-SLAM paper and closely related references mentioned in the source material. TWC-SLAM is a multi-agent cooperative LiDAR SLAM framework for similar indoor environments, particularly long corridors, repeated rooms, and repeated signs, in which pure geometric or point-cloud similarity is prone to aliasing. It integrates three information sources—LiDAR + IMU via FAST-LIO2, text semantics extracted from images, and WiFi fingerprint features derived from AP MACs and RSS values—to improve location identification, loop closure detection, and inter-agent map alignment. Its central mechanism is to use text semantics to propose candidate co-locations and WiFi features to validate whether those observations correspond to the same physical place, after which LiDAR registration is used to estimate relative transforms and produce a unified global map (Li et al., 26 Oct 2025).
1. Problem setting and design objective
Multi-agent cooperative SLAM in indoor buildings must solve two coupled problems: reliable loop closure for each agent and reliable location identification across agents. In similar indoor environments, these tasks are difficult because corridors, rooms, doors, and corners may share highly similar geometry and point distributions. Under those conditions, point cloud-based place recognition can add incorrect loop-closure or inter-agent constraints, which then warp the global pose graph and degrade the merged map.
The motivating failure modes in TWC-SLAM are explicitly multimodal rather than purely geometric. Using only text semantics can still be ambiguous when identical text appears at multiple positions, such as repeated “EXIT” signs or repeated room labels; a text observation may therefore be matched to a different physical location with the same string. Using only WiFi fingerprints is also unreliable, because RSS is noisy and nearby indoor positions can have similar AP compositions and signal patterns. TWC-SLAM is designed for the case in which LiDAR + IMU remains the geometric backbone, but text semantics and WiFi fingerprints are jointly used to disambiguate repetitive indoor structure (Li et al., 26 Oct 2025).
This design makes TWC-SLAM a cooperative indoor SLAM system specialized for semantic and radio-assisted place disambiguation rather than a generic replacement for LiDAR SLAM. A common misconception is to treat it as a text-based or WiFi-based SLAM pipeline. In fact, its architecture keeps LiDAR–IMU odometry and point-cloud registration central, while text and WiFi act as gating signals for robust co-location recognition.
2. System composition and sensing stack
The framework is organized into a multi-agent front-end odometry module, a text semantic matching module, a WiFi feature matching module, and a global mapping module. Each agent runs a single-agent front end based on FAST-LIO2, producing a local trajectory and a local point cloud map. Cameras capture environmental text, PaddleOCR extracts text regions and recognized strings, and WiFi receivers collect fingerprints composed of MAC address and RSS tuples. When text and WiFi jointly indicate a common location, the corresponding LiDAR frames are registered with ICP, and the resulting relative transforms are used to align sub-maps into a unified coordinate frame (Li et al., 26 Oct 2025).
| Component | Input | Role |
|---|---|---|
| Front-end odometry | LiDAR + IMU | Local trajectory and local point cloud map |
| Text semantic matching | Camera images | Candidate same-text pairs |
| WiFi feature matching | AP MACs + RSS | Candidate co-location validation |
| Global mapping | LiDAR frames / sub-maps | ICP alignment and map merging |
The sensing assumptions are explicit. Each agent is assumed to have a LiDAR, an IMU, a camera able to read text, and a WiFi receiver that is co-located with the LiDAR or connected by a known rigid transform. Clocks must be reasonably synchronized so that text and WiFi observations can be associated with LiDAR poses. The WiFi infrastructure is assumed to be relatively stable during mapping, and communication exists to share text/WiFi metadata and the necessary mapping information for cross-agent location recognition.
In the reported experiments, the geometric sensor is a Mid360 3D LiDAR, the WiFi receiver is a smartphone, and the platforms include a MetaCam EDU handheld device, a wheeled robot, and a legged robot. The environment consists of two floors of Guangming Laboratory in China, with repetitive room configurations, long corridors, duplicate text signs, and multiple WiFi APs. The agents typically observe more than eight distinct APs at a location.
3. Text semantics and WiFi fingerprints as location features
Text semantics in TWC-SLAM means human-readable environmental text used as semantic landmarks, such as room numbers and signs. PaddleOCR is used for scene text detection and recognition. Its loss is written as
with . For SLAM, each detected text instance is represented by its string , timestamp, and the agent pose at detection time. Cross-time and cross-agent text similarity is computed with a Levenshtein-based score
where is the edit distance. If , the pair is treated as a same-text candidate (Li et al., 26 Oct 2025).
WiFi features are built from AP name, MAC address, and RSS, but the AP identifier used for matching is the MAC address because SSID can map to multiple MACs. RSS is modeled as
although the method does not invert this expression to estimate distance. Instead, each location is associated with a WiFi fingerprint formed by denoising and averaging repeated RSS observations per MAC. After filtering samples whose deviation exceeds the estimated variability, the remaining values are averaged to obtain . A location fingerprint is thus represented as
The significance of these two feature types is complementary. Text provides sparse, human-readable labels with high semantic salience, but repeated signs induce aliasing. WiFi provides infrastructure-based local context, but with limited spatial resolution and susceptibility to multipath noise. TWC-SLAM treats neither modality as sufficient on its own; their joint use is intended specifically for repetitive indoor scenes.
4. Multi-modal location recognition and cooperative loop closure
The multi-modal location recognition algorithm is a hard-decision pipeline with thresholds , 0, and 1. It begins with text semantic gating: candidate pairs are generated only if the Levenshtein-based similarity satisfies 2. For those candidates, WiFi MAC overlap is computed as
3
where 4 and 5 are the numbers of MAC addresses at the two locations and 6 is the number of shared MAC addresses. If 7, the pair is rejected. For shared MACs, the RSS distance is then evaluated as
8
The pair is accepted as a same location only if the RSS comparison also passes thresholding (Li et al., 26 Oct 2025).
This procedure is used for both intra-agent loop closure and inter-agent location identification. For a newly observed text landmark with associated WiFi fingerprint and pose, the system compares it against the database of prior observations from the same agent and other agents. Text yields candidate correspondences, WiFi verifies physical co-location, and the accepted pairs become inputs to subsequent geometric registration.
An important detail is that the fusion is not probabilistic in the sense of a continuous joint score or learned multimodal embedding. The logic is an AND gate over text similarity, MAC overlap, and RSS distance. The source description also notes a minor inconsistency in the comparison with 9: the text states that the RSS distance “needs to be greater than the threshold 0,” while the conceptual interpretation is that smaller distance indicates greater similarity. The operational point, however, is clear: the RSS distance is thresholded to decide whether the two text observations plausibly originate from the same physical place.
The empirically selected thresholds for location recognition are 1, 2, and 3. This thresholded multimodal gating is the mechanism by which TWC-SLAM attempts to suppress false correspondences before any geometry-based alignment is attempted.
5. Geometric verification, map alignment, and global mapping
Once a same location has been confirmed by text and WiFi, TWC-SLAM performs geometric verification and relative pose estimation with ICP on the associated LiDAR frames or local sub-maps. If 4 denotes the reference cloud and 5 the source cloud, the transform is written as
6
and the ICP objective is
7
The resulting SE(3) transform becomes either an intra-agent loop-closure constraint or an inter-agent relative pose constraint (Li et al., 26 Oct 2025).
All sub-maps are then transformed into a common coordinate frame, and their union constitutes the global map. The paper describes this stage as “global optimization,” but it does not present an explicit factor-graph objective or a backend least-squares formulation such as a full pose-graph solver. It instead emphasizes FAST-LIO2 local maps together with pairwise ICP-based alignment once same locations have been recognized. This suggests a pose-graph interpretation, but the explicit optimization cost is not written in the source description.
The role of FAST-LIO2 is correspondingly delimited. It provides tightly coupled LiDAR–IMU odometry for each agent, with local consistency but possible drift. TWC-SLAM does not attempt to replace FAST-LIO2; it augments it with semantic and WiFi constraints so that drift can be corrected through loop closure and cross-agent alignment in environments where geometry alone is ambiguous.
6. Evaluation, limitations, and significance
The reported dataset consists of two scenes, Scene #01 and Scene #02, collected in Guangming Laboratory with a handheld device, a wheeled robot, and a legged robot. The main characteristics are repetitive rooms, long corridors, multiple identical text signs, and several WiFi APs. Evaluation uses same-location recognition precision and recall, together with End Point Error (EPE), defined as the distance between the trajectory starting point and endpoint when the start and end are supposed to coincide. Baselines and variants are DCL-SLAM, a text-only TWC-SLAM variant, a WiFi-only TWC-SLAM variant, and full TWC-SLAM (Li et al., 26 Oct 2025).
At the location-recognition level, the combined method achieves the best tradeoff at 8, 9: Scene #01 reports 0, 1, and Scene #02 reports 2, 3. The ablations show the expected failure modes. Text-only with low 4 gives high recall but low precision; increasing 5 improves precision to 78–88% but reduces recall to 31–45% because OCR variability causes mismatches. WiFi-only becomes more precise under stricter thresholds, up to about 87%, but recall drops in Scene #02 to about 40%.
The global-mapping results are summarized below.
| Scene | Travel distance (m) | EPE results |
|---|---|---|
| #01 | 265.32 | DCL-SLAM 1.69; text 0.35; WiFi 3.16; TWC-SLAM 0.21 |
| #02 | 463.91 | DCL-SLAM 1.41; text 1.67; WiFi 1.56; TWC-SLAM 0.16 |
These numbers reflect the environment-specific value of the multimodal design. In Scene #01, which contains four geometrically similar rooms, a long corridor, and repeated fire extinguisher cabinets with the same text, full TWC-SLAM reduces EPE to 0.21 m, compared with 1.69 m for DCL-SLAM, 0.35 m for the text-only variant, and 3.16 m for the WiFi-only variant. In Scene #02, which contains multiple similar rooms and repeated entry/exit text, TWC-SLAM reduces EPE to 0.16 m, compared with 1.41 m, 1.67 m, and 1.56 m for the same baselines and ablations. Qualitatively, the source description reports aligned trajectories, minimal ghosting, and crisp ceiling pipelines for TWC-SLAM, whereas the alternatives exhibit drift and point-cloud misalignment.
The limitations are also explicit. The method depends on the presence of usable text; in text-poor environments, its advantage diminishes. It assumes stable WiFi infrastructure, while dynamic AP changes or hotspots can distort fingerprints. Repeated identical text combined with similar WiFi fingerprints can still cause aliasing. WiFi RSS has limited spatial resolution, especially across neighboring rooms or corridors. The system is indoor-specific, and no clear benefit is claimed for outdoor or WiFi-poor settings. In addition, the absence of an explicitly described large-scale pose-graph backend may limit scalability in very large environments or with many agents.
Within cooperative indoor SLAM, the significance of TWC-SLAM lies in its treatment of repetitive architecture as a multimodal disambiguation problem rather than a purely geometric one. The authors present it as a multi-agent cooperative LiDAR SLAM framework that jointly uses text semantics and WiFi features for location identification and loop closure in similar indoor environments, and the reported experiments support the claim that neither text nor WiFi alone is sufficient in the target setting (Li et al., 26 Oct 2025).