Papers
Topics
Authors
Recent
Search
2000 character limit reached

Synesthesia of Machines: Sensing & Communication Integration

Updated 10 July 2026
  • Synesthesia of Machines is a paradigm that learns mappings between physical sensing and electromagnetic spaces using advanced neural methods.
  • It integrates diverse modalities such as LiDAR, RGB images, and pilot signals to enhance channel modeling, precoding, and collaborative sensing.
  • The approach leverages domain priors, latent-space alignment, and task-aware encoding to improve performance in wireless transmission and environment inference.

Searching arXiv for recent SoM-related papers to ground the article with up-to-date references. Using arXiv search for "Synesthesia of Machines" and related wireless sensing-communication papers. Synesthesia of Machines (SoM) is a research paradigm inspired by human synesthesia and framed in the current literature as intelligent multi-modal sensing-communication integration. Across recent arXiv work, SoM denotes learned mappings between physical sensing spaces and electromagnetic or communication spaces, or the joint processing of semantic data and physical-layer side information, so that a machine can infer scatterers, pathloss, multipath, CSI, beam policies, or transmission strategies from sensory observations and channel context rather than treating communication as a standalone noisy pipe (Li et al., 14 Sep 2025, Huang et al., 2024, Li et al., 2 Sep 2025).

1. Conceptual definition and scope

The strongest common definition in the literature is cross-modal mapping. In vehicular channel modeling, SoM is explicitly the “mapping relationship between physical environment and electromagnetic space,” with LiDAR point clouds representing physical environment space and scatterers representing electromagnetic space (Huang et al., 2024). In wireless image transmission, SoM is formulated as using artificial neural networks “to leverage artificial neural networks to extract high-density, task-aware, and robust features, thereby achieving intelligent integration of communication and multi-modal sensing,” then specialized to the tight coupling of image processing and physical-layer wireless transmission through side information such as SNR, time-varying CSI, Doppler, and channel aging (Li et al., 14 Sep 2025). In task-driven MIMO transmission, SoM is described as the “intelligent integration of multi-modal information from disparate ‘machine senses’, such as communication devices and sensors,” realized as joint encoding of perception features and channel information into a task-oriented representation (Li et al., 2 Sep 2025).

A second recurring definition is latent-space alignment. In WiCo-MG and WiCo-PG, SoM is instantiated as paired cross-modal feature spaces, with a sensing SoM space and a channel SoM space connected by a learned mapping; in WiCo-MG these spaces are discrete-continuous for sensing images and discrete for multipath maps, while WiCo-PG uses dual VQGANs for RGB images and pathloss maps (Han et al., 19 Nov 2025, Sun et al., 19 Nov 2025). In LLM-based variants, SoM also includes alignment between sensory features and the semantic space of a pretrained LLM: GPT-2 in LLM4PG and LLaMA 3.2 in LLM4MG (Sun et al., 4 Nov 2025, Huang et al., 18 Sep 2025).

The scope of SoM in the current corpus is broad. It covers wireless image transmission, task-driven MIMO, LiDAR point cloud transmission, scatterer recognition, pathloss map generation, multipath generation, CSI learning, FDD precoding, online precoding under sensing heterogeneity, and sub-THz ISAC for air-ground links (Li et al., 14 Sep 2025, Li et al., 2 Sep 2025, Liu et al., 8 Sep 2025, Zhang et al., 2024, Zhang et al., 9 Jun 2025, Yang et al., 15 Jun 2025). This suggests that SoM is best understood not as a single algorithm, but as a design doctrine for coupling sensing and communication through learned, physically informed representations.

2. Recurrent technical structure

A consistent technical pattern is the insertion of sensing or physical-layer variables directly into intermediate neural processing. In DCAT for wireless image transmission, the side-information set is

P={SNR,H^,fD,τag},\bm{\mathcal{P}}=\Big\{\text{SNR},\hat{\bm{H}},\overline{f_D},\bm{\tau_{ag}}\Big\},

and the encoder and decoder are explicitly functions of both the source image and this physical-layer information (Li et al., 14 Sep 2025). In LE-CLN for wideband multi-user CSI learning, raw LiDAR is converted into RF-oriented features including a range map, an equivalent small-scale fading map that labels receiver and nearby-vehicle scatterers, and an equivalent large-scale fading map, after which adaptive feature weighting controls how much the network trusts LiDAR versus pilots (Zhang et al., 2024). In sub-THz ISAC, SoM ties visual localization, squint-aware beam management, and hybrid precoding to a communication-sensing channel correlation metric, with true-time delay and phase-shifter degrees of freedom treated as learnable control variables rather than fixed hardware details (Yang et al., 15 Jun 2025).

A second pattern is the explicit use of domain priors and interpretable regularizers. In scatterer recognition for vehicular channel modeling, visibility regions enforce propagation-aware filtering after MLP-based LiDAR-to-scatterer mapping (Huang et al., 2024). In DCAT, the attention path is regularized by CSI NMSE and the permutation path by actual symbol impairment, linking learned behavior to interpretable physical quantities (Li et al., 14 Sep 2025). In online FDD precoding, the pseudo-CSI simulator uses semi-supervised learning and an online loss based on observed pilots, rather than assuming access to ground-truth downlink CSI labels (Zhang et al., 9 Jun 2025).

A third pattern is modular latent alignment. WiCo-MG aligns sensing images and multipath maps through sensing-initialized codebooks, discrete-continuous SoM spaces, and a frequency-aware shared-routed mixture of experts Transformer (Han et al., 19 Nov 2025). WiCo-PG uses dual VQGANs and a frequency-guided shared-routed MoE Transformer to map RGB images to pathloss maps (Sun et al., 19 Nov 2025). LLM4PG and LLM4MG instead align multi-modal sensing features with a pretrained language-model latent space, treating the LLM as a cross-domain reasoning engine rather than a text generator (Sun et al., 4 Nov 2025, Huang et al., 18 Sep 2025). A plausible implication is that SoM has developed a recognizable design grammar: modality-specific encoders, explicit physical priors, latent alignment, and task-aware decoders.

3. End-to-end semantic transmission

In wireless transmission problems, SoM is used to redesign what is transmitted, not merely how it is protected. DCAT, the SoM realization in wireless image transmission over complex dynamic channels, couples a Swin Transformer backbone with Dynamic Channel-Attention and Dynamic Channel-Permutation modules. The former performs global feature-channel reweighting conditioned on SNR, CSI, Doppler, and aging; the latter performs index-level arrangement of real and imaginary feature channels in time according to predicted reliability. Under the “Aging Scenario,” DCAT reports an average PSNR gain of approximately $0.431$ dB over the best baseline, a 10.5%10.5\% improvement, and an LPIPS reduction of approximately 12.6%12.6\%, while gains increase from $0.089$ dB at CR=1/12\text{CR}=1/12 to $0.936$ dB at CR=1/3\text{CR}=1/3 (Li et al., 14 Sep 2025).

SoM-MIMO applies the same principle to task-driven digital MIMO. Instead of transmitting images for later reconstruction, it compresses the Mask R-CNN feature pyramid through hierarchical feature fusion, adapts the compact feature using MIMO channel-aware encoding and decoding, and optimizes directly for instance segmentation. At CR=1/16\text{CR}=1/16, N=2N=2, and $0.431$0, it reports average mAP improvements of $0.431$1 points over Swin-JSCC and $0.431$2 points over ViT-JSCC-MIMO across all SNR levels, with larger gains of $0.431$3 and $0.431$4 points at $0.431$5 dB (Li et al., 2 Sep 2025).

LPC-FT extends SoM to collaborative LiDAR point cloud transmission. It combines density-preserving deep point cloud compression, a self-attention channel encoder, a cross-attention fusion module that integrates transmitted and local receiver features, and a nonlinear activation layer inserted before digital quantization and modulation. Against traditional octree-based compression followed by channel coding, and against prior deep compression or semantic communication baselines, LPC-FT reports an average Chamfer Distance reduction of $0.431$6 and an average PSNR improvement of $0.431$7 dB (Liu et al., 8 Sep 2025).

These systems share a common shift in abstraction. They no longer optimize a transparent bit pipe around raw signals; they optimize a semantic latent shaped jointly by sensing structure, channel behavior, and downstream objectives. This suggests that, in transmission-oriented SoM, the primary object of design is the learned representation itself.

4. Environment-to-channel generation and channel modeling

A major SoM line of work treats channel generation as a cross-modal inference problem. In “Scatterer Recognition for Multi-Modal Intelligent Vehicular Channel Modeling via Synesthesia of Machines,” LiDAR point clouds are fused, clustered, co-registered with ray-tracing outputs, and mapped by an MLP to scatterer numbers per object. Scatterer recognition exceeds $0.431$8 accuracy across all vehicular traffic densities and streets, achieves an average accuracy of $0.431$9, improves accuracy by more than 10.5%10.5\%0 over a random-generation baseline, and yields power delay profiles that closely match ray-tracing results (Huang et al., 2024).

The V2V multi-modal intelligent channel model based on ScaR replaces object-level scatterer counting with a SegNet-based grid-map predictor. LiDAR density, height, and distance grids are mapped to scatterer density grids, after which scatterers are classified as static, dynamic, or unknown and used to drive a non-stationary channel model. Reported scatterer-grid classification accuracy exceeds 10.5%10.5\%1, density accuracy is 10.5%10.5\%2 in high vehicular traffic density and 10.5%10.5\%3 to 10.5%10.5\%4 in medium and low density, and the resulting TF-CF, TACF, and DPSD match ray-tracing well (Han et al., 13 Jan 2025). LLM4SG moves the same LiDAR-to-scatterer problem into a truncated GPT-2 backbone with positional encoding and frequency embedding, reporting 10.5%10.5\%5, 10.5%10.5\%6, and TACF behavior close to ray tracing in downstream channel modeling (Han et al., 23 May 2025).

Pathloss and multipath generation push SoM toward foundation-model territory. LLM4PG adapts GPT-2 for RGB-D-plus-frequency to pathloss map generation and reports a best NMSE of 10.5%10.5\%7, more than 10.5%10.5\%8 dB better than a conventional GAN baseline, with cross-condition generalization NMSE of 10.5%10.5\%9 and a 12.6%12.6\%0 dB gain over baseline (Sun et al., 4 Nov 2025). WiCo-PG replaces the language-model backbone with dual VQGANs plus a frequency-guided shared-routed MoE Transformer, reaches NMSE 12.6%12.6\%1, exceeds LLM4PG and a conventional deep model by more than 12.6%12.6\%2 dB, and outperforms LLM4PG by at least 12.6%12.6\%3 dB using 12.6%12.6\%4 samples in few-shot generalization (Sun et al., 19 Nov 2025). WiCo-MG extends the same template from pathloss to multipath-map generation and reports more than 12.6%12.6\%5 dB NMSE reduction over baselines, together with improved out-of-distribution generalization and scaling with both model and dataset size (Han et al., 19 Nov 2025). LLM4MG adapts LLaMA 3.2 to multipath generation from mmWave radar, RGB-D, and LiDAR, achieving LoS/NLoS classification accuracy of 12.6%12.6\%6, power-generation NMSE of 12.6%12.6\%7, delay-generation NMSE of 12.6%12.6\%8, and validated real-world generalization (Huang et al., 18 Sep 2025).

Across these works, SoM changes channel modeling from parameter fitting on RF observations alone to a generative mapping from sensed environments to electromagnetic structure. A plausible implication is that SoM is becoming a bridge between wireless digital twins, channel foundation models, and environment-aware simulation.

5. CSI learning, precoding, and ISAC control

SoM is also used for control problems in which sensing substitutes for explicit channel acquisition. LE-CLN for wideband multi-user MIMO-OFDM CSI learning transforms LiDAR into RF-relevant features and combines them with user-localized over-complete angular pilot measurements. Its adaptive feature weight control module learns when LiDAR should dominate and when pilots should dominate; the reported outcome is higher channel-estimation accuracy and spectrum efficiency than OMP, AMP, LS, CENN, and GM-LAMP, especially in latency-sensitive settings where pilot transmissions are reduced (Zhang et al., 2024).

In FDD precoding, SoM is coupled with vertical federated learning to handle heterogeneous on-board sensing. The offline VFL scheme maps each vehicle’s short pilot observation and available sensing features—GPS, RGB, LiDAR—to its own precoding vector, and numerical results show that short-pilot SoM-aided precoding closely approximates traditional optimization with perfect CSI (Zhang et al., 19 Jan 2025). The online extension adds a pseudo downlink CSI label simulator trained with semi-supervised learning and an online loss, enabling label-free adaptation when user number or sensing configuration changes. That system reports a 12.6%12.6\%9 reduction in pilot sequence length while remaining close to traditional optimization with perfect CSI (Zhang et al., 9 Jun 2025).

In sub-THz air-ground ISAC, SoM integrates RGB-D cameras, squint-aware beam management, and a learned hybrid precoder. The framework defines a communication-sensing channel correlation

$0.089$0

uses visual priors to localize users and targets, applies a squint-aware cross-pattern beam-tracking procedure to refine 3D angles, and then uses ViR-Net to predict true-time-delay settings, phase-shifter phases, and digital precoders (Yang et al., 15 Jun 2025). This places SoM beyond inference and into hardware-constrained control.

The common structure in these papers is that sensing is not merely auxiliary metadata. It becomes an operational surrogate for pilots, a prior on beamspace structure, or a direct input to the control law. This suggests that SoM extends naturally from perception and modeling into closed-loop wireless optimization.

6. Datasets, evidence regimes, and limitations

The evidence base for SoM is unusually dataset-centric. Besides task-specific corpora such as UDIS-D for wireless image transmission, Cityscapes for task-driven MIMO, and OpV2V for LiDAR point cloud transmission, SoM research repeatedly constructs aligned sensing-communication datasets: V2V-M3 for LiDAR-to-scatterer learning, SynthSoM-U2G for RGB-D-to-pathloss generation, and SynthSoM-Twin for Sim2Real transfer (Li et al., 14 Sep 2025, Li et al., 2 Sep 2025, Liu et al., 8 Sep 2025, Han et al., 23 May 2025, Sun et al., 4 Nov 2025, Chen et al., 14 Nov 2025). SynthSoM-Twin is especially notable because it imports spatio-temporally consistent static and dynamic scenes into AirSim, WaveFarer, and Sionna RT, produces $0.089$1 synthetic snapshots across RGB, depth, LiDAR, mmWave radar, pathloss, and small-scale fading, and shows that injecting less than $0.089$2 of real-world data can achieve similar and even better downstream performance than training with all real-world data only (Chen et al., 14 Nov 2025).

At the same time, limitations are recurrent. Several SoM channel-generation papers are trained entirely on synthetic data and explicitly note potential domain gaps to real measurements, simplified material models, limited scenario coverage, or restricted sensing modalities (Han et al., 19 Nov 2025, Sun et al., 4 Nov 2025). The DCAT image-transmission system is currently single-user, single-carrier, and SISO, with future directions pointing to MIMO, OFDM, multi-user interference, and hardware tests (Li et al., 14 Sep 2025). LLM4PG focuses on pathloss rather than small-scale fading, delay spread, or MIMO matrices (Sun et al., 4 Nov 2025). SynthSoM-Twin itself remains scenario-specific and relies on reconstructed dynamic objects and simulator fidelity (Chen et al., 14 Nov 2025).

A frequent misconception is to equate SoM with generic sensor fusion. The corpus does not support that reduction. In most formulations, SoM denotes either a learned mapping from one physical domain to another, or a communication architecture in which physical-layer and sensing information explicitly modulate representation formation, decoding, or control (Huang et al., 2024, Li et al., 14 Sep 2025, Li et al., 2 Sep 2025). Another misconception is that SoM implies a single preferred model class. In practice, SoM has been implemented with MLPs, SegNet, Swin Transformer, ResNet-FPN, GPT-2, LLaMA, VQGANs, mixture-of-experts Transformers, and vertical federated learning (Han et al., 13 Jan 2025, Sun et al., 4 Nov 2025, Sun et al., 19 Nov 2025, Zhang et al., 9 Jun 2025). What unifies these systems is not architecture but the decision to make sensing and communication co-inform each other in a physically meaningful way.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (15)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Synesthesia of Machines (SoM).