Papers
Topics
Authors
Recent
Search
2000 character limit reached

TWC-SLAM: Multimodal Cooperative LiDAR SLAM

Updated 4 July 2026
  • TWC-SLAM is a multi-agent cooperative LiDAR SLAM framework that fuses LiDAR+IMU data with text semantics and WiFi fingerprints to resolve aliasing in similar indoor layouts.
  • It employs a multimodal gating strategy where text similarity and WiFi fingerprint matching generate candidate co-locations that are then verified using LiDAR registration via ICP.
  • Experiments in repetitive indoor settings demonstrate that TWC-SLAM reduces endpoint errors to 0.16–0.21 m, outperforming single-modality approaches in loop closure and map alignment.

Searching arXiv for the TWC-SLAM paper and closely related references mentioned in the source material. TWC-SLAM is a multi-agent cooperative LiDAR SLAM framework for similar indoor environments, particularly long corridors, repeated rooms, and repeated signs, in which pure geometric or point-cloud similarity is prone to aliasing. It integrates three information sources—LiDAR + IMU via FAST-LIO2, text semantics extracted from images, and WiFi fingerprint features derived from AP MACs and RSS values—to improve location identification, loop closure detection, and inter-agent map alignment. Its central mechanism is to use text semantics to propose candidate co-locations and WiFi features to validate whether those observations correspond to the same physical place, after which LiDAR registration is used to estimate relative transforms and produce a unified global map (Li et al., 26 Oct 2025).

1. Problem setting and design objective

Multi-agent cooperative SLAM in indoor buildings must solve two coupled problems: reliable loop closure for each agent and reliable location identification across agents. In similar indoor environments, these tasks are difficult because corridors, rooms, doors, and corners may share highly similar geometry and point distributions. Under those conditions, point cloud-based place recognition can add incorrect loop-closure or inter-agent constraints, which then warp the global pose graph and degrade the merged map.

The motivating failure modes in TWC-SLAM are explicitly multimodal rather than purely geometric. Using only text semantics can still be ambiguous when identical text appears at multiple positions, such as repeated “EXIT” signs or repeated room labels; a text observation may therefore be matched to a different physical location with the same string. Using only WiFi fingerprints is also unreliable, because RSS is noisy and nearby indoor positions can have similar AP compositions and signal patterns. TWC-SLAM is designed for the case in which LiDAR + IMU remains the geometric backbone, but text semantics and WiFi fingerprints are jointly used to disambiguate repetitive indoor structure (Li et al., 26 Oct 2025).

This design makes TWC-SLAM a cooperative indoor SLAM system specialized for semantic and radio-assisted place disambiguation rather than a generic replacement for LiDAR SLAM. A common misconception is to treat it as a text-based or WiFi-based SLAM pipeline. In fact, its architecture keeps LiDAR–IMU odometry and point-cloud registration central, while text and WiFi act as gating signals for robust co-location recognition.

2. System composition and sensing stack

The framework is organized into a multi-agent front-end odometry module, a text semantic matching module, a WiFi feature matching module, and a global mapping module. Each agent runs a single-agent front end based on FAST-LIO2, producing a local trajectory and a local point cloud map. Cameras capture environmental text, PaddleOCR extracts text regions and recognized strings, and WiFi receivers collect fingerprints composed of MAC address and RSS tuples. When text and WiFi jointly indicate a common location, the corresponding LiDAR frames are registered with ICP, and the resulting relative transforms are used to align sub-maps into a unified coordinate frame (Li et al., 26 Oct 2025).

Component Input Role
Front-end odometry LiDAR + IMU Local trajectory and local point cloud map
Text semantic matching Camera images Candidate same-text pairs
WiFi feature matching AP MACs + RSS Candidate co-location validation
Global mapping LiDAR frames / sub-maps ICP alignment and map merging

The sensing assumptions are explicit. Each agent is assumed to have a LiDAR, an IMU, a camera able to read text, and a WiFi receiver that is co-located with the LiDAR or connected by a known rigid transform. Clocks must be reasonably synchronized so that text and WiFi observations can be associated with LiDAR poses. The WiFi infrastructure is assumed to be relatively stable during mapping, and communication exists to share text/WiFi metadata and the necessary mapping information for cross-agent location recognition.

In the reported experiments, the geometric sensor is a Mid360 3D LiDAR, the WiFi receiver is a smartphone, and the platforms include a MetaCam EDU handheld device, a wheeled robot, and a legged robot. The environment consists of two floors of Guangming Laboratory in China, with repetitive room configurations, long corridors, duplicate text signs, and multiple WiFi APs. The agents typically observe more than eight distinct APs at a location.

3. Text semantics and WiFi fingerprints as location features

Text semantics in TWC-SLAM means human-readable environmental text used as semantic landmarks, such as room numbers and signs. PaddleOCR is used for scene text detection and recognition. Its loss is written as

L=Ls+λgLg,L = L_s + \lambda_g L_g,

with λg=1\lambda_g = 1. For SLAM, each detected text instance is represented by its string SS, timestamp, and the agent pose at detection time. Cross-time and cross-agent text similarity is computed with a Levenshtein-based score

Sij=max⁡(∣Si∣,∣Sj∣)−d(Si,Sj)max⁡(∣Si∣,∣Sj∣),S_{ij} = \frac{\max(|S_i|, |S_j|) - d(S_i, S_j)}{\max(|S_i|, |S_j|)},

where d(Si,Sj)d(S_i,S_j) is the edit distance. If Sij≥αS_{ij} \ge \alpha, the pair is treated as a same-text candidate (Li et al., 26 Oct 2025).

WiFi features are built from AP name, MAC address, and RSS, but the AP identifier used for matching is the MAC address because SSID can map to multiple MACs. RSS is modeled as

RSS=Pt−K−10ζlog⁡10d,RSS = P_t - K - 10 \zeta \log_{10} d,

although the method does not invert this expression to estimate distance. Instead, each location is associated with a WiFi fingerprint formed by denoising and averaging repeated RSS observations per MAC. After filtering samples whose deviation exceeds the estimated variability, the remaining values are averaged to obtain RSS∗RSS^*. A location fingerprint is thus represented as

L:{(MAC1:RSS1∗),(MAC2:RSS2∗),…,(MACn:RSSn∗)}.L : \{(MAC_1 : RSS_1^*), (MAC_2 : RSS_2^*), \dots, (MAC_n : RSS_n^*)\}.

The significance of these two feature types is complementary. Text provides sparse, human-readable labels with high semantic salience, but repeated signs induce aliasing. WiFi provides infrastructure-based local context, but with limited spatial resolution and susceptibility to multipath noise. TWC-SLAM treats neither modality as sufficient on its own; their joint use is intended specifically for repetitive indoor scenes.

4. Multi-modal location recognition and cooperative loop closure

The multi-modal location recognition algorithm is a hard-decision pipeline with thresholds α\alpha, λg=1\lambda_g = 10, and λg=1\lambda_g = 11. It begins with text semantic gating: candidate pairs are generated only if the Levenshtein-based similarity satisfies λg=1\lambda_g = 12. For those candidates, WiFi MAC overlap is computed as

λg=1\lambda_g = 13

where λg=1\lambda_g = 14 and λg=1\lambda_g = 15 are the numbers of MAC addresses at the two locations and λg=1\lambda_g = 16 is the number of shared MAC addresses. If λg=1\lambda_g = 17, the pair is rejected. For shared MACs, the RSS distance is then evaluated as

λg=1\lambda_g = 18

The pair is accepted as a same location only if the RSS comparison also passes thresholding (Li et al., 26 Oct 2025).

This procedure is used for both intra-agent loop closure and inter-agent location identification. For a newly observed text landmark with associated WiFi fingerprint and pose, the system compares it against the database of prior observations from the same agent and other agents. Text yields candidate correspondences, WiFi verifies physical co-location, and the accepted pairs become inputs to subsequent geometric registration.

An important detail is that the fusion is not probabilistic in the sense of a continuous joint score or learned multimodal embedding. The logic is an AND gate over text similarity, MAC overlap, and RSS distance. The source description also notes a minor inconsistency in the comparison with λg=1\lambda_g = 19: the text states that the RSS distance “needs to be greater than the threshold SS0,” while the conceptual interpretation is that smaller distance indicates greater similarity. The operational point, however, is clear: the RSS distance is thresholded to decide whether the two text observations plausibly originate from the same physical place.

The empirically selected thresholds for location recognition are SS1, SS2, and SS3. This thresholded multimodal gating is the mechanism by which TWC-SLAM attempts to suppress false correspondences before any geometry-based alignment is attempted.

5. Geometric verification, map alignment, and global mapping

Once a same location has been confirmed by text and WiFi, TWC-SLAM performs geometric verification and relative pose estimation with ICP on the associated LiDAR frames or local sub-maps. If SS4 denotes the reference cloud and SS5 the source cloud, the transform is written as

SS6

and the ICP objective is

SS7

The resulting SE(3) transform becomes either an intra-agent loop-closure constraint or an inter-agent relative pose constraint (Li et al., 26 Oct 2025).

All sub-maps are then transformed into a common coordinate frame, and their union constitutes the global map. The paper describes this stage as “global optimization,” but it does not present an explicit factor-graph objective or a backend least-squares formulation such as a full pose-graph solver. It instead emphasizes FAST-LIO2 local maps together with pairwise ICP-based alignment once same locations have been recognized. This suggests a pose-graph interpretation, but the explicit optimization cost is not written in the source description.

The role of FAST-LIO2 is correspondingly delimited. It provides tightly coupled LiDAR–IMU odometry for each agent, with local consistency but possible drift. TWC-SLAM does not attempt to replace FAST-LIO2; it augments it with semantic and WiFi constraints so that drift can be corrected through loop closure and cross-agent alignment in environments where geometry alone is ambiguous.

6. Evaluation, limitations, and significance

The reported dataset consists of two scenes, Scene #01 and Scene #02, collected in Guangming Laboratory with a handheld device, a wheeled robot, and a legged robot. The main characteristics are repetitive rooms, long corridors, multiple identical text signs, and several WiFi APs. Evaluation uses same-location recognition precision and recall, together with End Point Error (EPE), defined as the distance between the trajectory starting point and endpoint when the start and end are supposed to coincide. Baselines and variants are DCL-SLAM, a text-only TWC-SLAM variant, a WiFi-only TWC-SLAM variant, and full TWC-SLAM (Li et al., 26 Oct 2025).

At the location-recognition level, the combined method achieves the best tradeoff at SS8, SS9: Scene #01 reports Sij=max⁡(∣Si∣,∣Sj∣)−d(Si,Sj)max⁡(∣Si∣,∣Sj∣),S_{ij} = \frac{\max(|S_i|, |S_j|) - d(S_i, S_j)}{\max(|S_i|, |S_j|)},0, Sij=max⁡(∣Si∣,∣Sj∣)−d(Si,Sj)max⁡(∣Si∣,∣Sj∣),S_{ij} = \frac{\max(|S_i|, |S_j|) - d(S_i, S_j)}{\max(|S_i|, |S_j|)},1, and Scene #02 reports Sij=max⁡(∣Si∣,∣Sj∣)−d(Si,Sj)max⁡(∣Si∣,∣Sj∣),S_{ij} = \frac{\max(|S_i|, |S_j|) - d(S_i, S_j)}{\max(|S_i|, |S_j|)},2, Sij=max⁡(∣Si∣,∣Sj∣)−d(Si,Sj)max⁡(∣Si∣,∣Sj∣),S_{ij} = \frac{\max(|S_i|, |S_j|) - d(S_i, S_j)}{\max(|S_i|, |S_j|)},3. The ablations show the expected failure modes. Text-only with low Sij=max⁡(∣Si∣,∣Sj∣)−d(Si,Sj)max⁡(∣Si∣,∣Sj∣),S_{ij} = \frac{\max(|S_i|, |S_j|) - d(S_i, S_j)}{\max(|S_i|, |S_j|)},4 gives high recall but low precision; increasing Sij=max⁡(∣Si∣,∣Sj∣)−d(Si,Sj)max⁡(∣Si∣,∣Sj∣),S_{ij} = \frac{\max(|S_i|, |S_j|) - d(S_i, S_j)}{\max(|S_i|, |S_j|)},5 improves precision to 78–88% but reduces recall to 31–45% because OCR variability causes mismatches. WiFi-only becomes more precise under stricter thresholds, up to about 87%, but recall drops in Scene #02 to about 40%.

The global-mapping results are summarized below.

Scene Travel distance (m) EPE results
#01 265.32 DCL-SLAM 1.69; text 0.35; WiFi 3.16; TWC-SLAM 0.21
#02 463.91 DCL-SLAM 1.41; text 1.67; WiFi 1.56; TWC-SLAM 0.16

These numbers reflect the environment-specific value of the multimodal design. In Scene #01, which contains four geometrically similar rooms, a long corridor, and repeated fire extinguisher cabinets with the same text, full TWC-SLAM reduces EPE to 0.21 m, compared with 1.69 m for DCL-SLAM, 0.35 m for the text-only variant, and 3.16 m for the WiFi-only variant. In Scene #02, which contains multiple similar rooms and repeated entry/exit text, TWC-SLAM reduces EPE to 0.16 m, compared with 1.41 m, 1.67 m, and 1.56 m for the same baselines and ablations. Qualitatively, the source description reports aligned trajectories, minimal ghosting, and crisp ceiling pipelines for TWC-SLAM, whereas the alternatives exhibit drift and point-cloud misalignment.

The limitations are also explicit. The method depends on the presence of usable text; in text-poor environments, its advantage diminishes. It assumes stable WiFi infrastructure, while dynamic AP changes or hotspots can distort fingerprints. Repeated identical text combined with similar WiFi fingerprints can still cause aliasing. WiFi RSS has limited spatial resolution, especially across neighboring rooms or corridors. The system is indoor-specific, and no clear benefit is claimed for outdoor or WiFi-poor settings. In addition, the absence of an explicitly described large-scale pose-graph backend may limit scalability in very large environments or with many agents.

Within cooperative indoor SLAM, the significance of TWC-SLAM lies in its treatment of repetitive architecture as a multimodal disambiguation problem rather than a purely geometric one. The authors present it as a multi-agent cooperative LiDAR SLAM framework that jointly uses text semantics and WiFi features for location identification and loop closure in similar indoor environments, and the reported experiments support the claim that neither text nor WiFi alone is sufficient in the target setting (Li et al., 26 Oct 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to TWC-SLAM.